[0] https://dank.systems/posts/2026-09-15-ai-bear.html
[1] https://gowers.wordpress.com/2026/08/12/what-sort-of-maths-a...
[0] https://dank.systems/posts/2026-09-15-ai-bear.html
[1] https://gowers.wordpress.com/2026/08/12/what-sort-of-maths-a...
But they’re already extending into politics, military, journalism, art, and many other fields that aren’t verifiable in any meaningful sense of the word.
What makes you say that? What is an example of a domain where the improvement is small?
I can't think of any at all. Compare something as unverifiable as "Make good music". Models now are many times better than 3 years ago.
Improvement in this context means "better quality results".
You can use better quality models to do worse things with.
I'm not making any claim about second order effects like that.
I'm not aware of any benchmarks that measure the first "analyze XYZ geopolitical situation" but legal reasoning is very closely related to "explain the ramifications of XYZ law".
Legal Bench[1] measures legal reasoning. It's close to saturated (ie, there isn't a lot of room for improvement) but Fable scores 88% vs eg Opus 4.7 at 85%.
There probably isn't a lot of room for improvement on something like this - there is enough disagreement in legal reasoning to mean 100% is going to be impossible.
These models are spikey as hell, they can gain incredible capabilities but that doesn't make them godlike minds, it's more like a scaled up version of rain man.
I think there is a roof on the level of reasoning needed for legal reasoning, and I think we are pretty close to it. Legal arguments just aren't that complex in terms of logic chains.
I'd bet the cyclometric complexity of any given court case is lower than a dense piece of code for example.
Just realized this is closely related to geo-political forecasting, where AI is now beating the best humans
https://www.economist.com/science-and-technology/2026/09/16/...
We'll have plenty of time for this, while living off UBI.
Really? Do better.