10,961 karma · joined April 24, 2021
TypeScript has different semantics depending on settings in tsconfig e.g. useDefineForClassFields
Anthropic is just trying to get to ASI and relying on the blind faith of investors that there will be a return at the end.
What matters is which strategy is more effective for attracting investor capital, because neither can survive without it.
It seems that 6.1 Sol is closely related to Astra. As for whatever 6.0 Sol was, it's anyone's guess. I'd speculate that 6.0 was just 5.6 with post training. It's possible however that it is the same model but with severe performance issues resolved so they bumped the version number.
There may have been some human involvement with this website but they certainly didn’t edit it for style. Unfortunately if you don’t edit for style then people will default to the assumption that it’s entirely vibecoded.
An example of allophony in English is the aspirated /pʰ/ in pin and the unaspirated /p/ in spin, which to a speaker of Hindi and many other languages would be considered two distinct sounds.
- Disproof of the Jacobian conjecture by example
- Construction of a non-sofic group
- Existence of singularity in Navier-Stokes
Mathematical conjectures tend to be universally quantified, especially those conjectures that are used as building blocks (e.g. RH). If anything, AI models are currently performing a useful service by disproving false conjectures, a kind of mathematical weeding.
The good news from the last couple of years of coding agents is that while models have become more persistent and knowledgable, their creativity (defined as being able to escape their training distribution and synthesize completely novel ideas) is improving at a much slower rate.
AI will only become a threat to mathematics if/when it develops the capability for creative big-picture problem solving. If that happens, the impact on mathematics will be a footnote compared to the impacts on society at large, since creativity unlocks a host of new economic capabilities.
That they heard a rumour that a major open problem had been solved, so they decided to try and scoop the other mathematicians while they were writing up their preprint is extremely unsporting.
Then they decided to exclude an author because of his employer, even though he had used their own products to write the proof!
They haven't necessarily breached any formal ethical rules but their behaviour will lead to them and their products being shut out from the mathematical community.
In chess, depth usually wins because of how narrow the search tree is compared e.g. to Go.
Were it not for copyright then BSDs could take code from Linux and perhaps there'd be less of a monoculture, for example.
What exactly would AI have to do in order to not be called a bubble?
Although I believe there's a difference with the Manhattan project. The original atom bomb was "useful" even at its small scale. Whereas there's likely to be a lot of time between a quantum computer that can factor 69 and a quantum computer that can factor RSA 2048.
This is one of the only benchmarks that actually matters for testing the frontier however. Other benchmarks can be gamed by simply being more persistent, but HLE is a diverse set of open-ended research-level questions. It tests domain knowledge and problem solving skills. Burning more reasoning tokens may help somewhat but not as much as e.g. coding benchmarks.