anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.
(I'm not happy about the above being true, but it's the reality I seem to inhabit.)
Benchmarks are useful but only on a log2 basis. One model performing at 50% and another at 75% is just as impressive as one model performing at 78% and another at 90%. Confoundingly, a benchmark becomes useless once a frontier model scores over ~95% on them.
This technology is strictly an extractive parasite on the world. Use it, but don't be excited.
i think that there is growing organized labor today that produces no surplus. instead, it transfers wealth from some to others, causing net harm to all in the process. an example of this would be purdue pharma.
depending on who you ask the list of jobs and industries which have zero surplus is getting large. swathes of private equity and leveraged financial instruments, shitcoins, management consultancy, are pure deadweight loss.
the work does nothing or causes net harm.
That's the opposite of parasitic.
People already started using contributor API, and your input is irrelevant.
Compare that to Muse spark 1.3
$1.25/M input, $4.25/M output (without data sharing) $0.10/M input, $0.20/M output (with data sharing)
It is dirt cheap, but only if you are willing to share your data with meta and allow them to use it for improving their models and products.
This is an error I would expect from sonnet 4, not a model that was supposedly just a few points behind sol.