Opus historically had issues with minor typographical errors, though recently that seems to not happen often, lots of very sharp people at Anthropic.
So a month ago if I wanted something from Opus I’d run it through a cleanup pass courtesy of one of the other ones, but even my old standby dolphin-8x7 can clean up typos. 1106 can as well, but all else equal I don’t want to be sending my stuff to any black box data warehouse and I’m always surprised so many other sophisticated people don’t share the preference.
My personal eyeball capability check is to posit a gauge symmetry and ask what it thinks the implied conserved quantity is, and I’ve yet to see Opus not crush that relative to anything else, including real footnotes.
On coding I usually hand it a Prisma schema and ask for a proto3/gRPC definition that is a good way to interact with it, Opus in my personal experience also dominates there.
If you have an example of a task that represents a counter example I’d be grateful for another little integration test for my personal ad-hoc model card. I want to know the best tool for every job.
I really don't like how condescending your root comment is, I don't have any idea how all these extra things you're disappointed in are relevant at all to the actual topic.
I’d likewise be grateful if you’d lend me an example or two that should be on my little ad-hoc task set.
Assuming the former, here's an example I had yesterday where Opus was never able to formulate a correct response even with multiple follow-ups but GPT-4o got it with only two follow ups:
> I have a postgres database with a table called projects. On that table is a column called designs that is of json type and is nullable as well as a column called id that is of type text. The data in the designs column will be an array of objects (assuming it's not null). Each object in the array will have a field called price that will be a string. Most of these prices will be in the format $#,### where # is a number. However some will be in the format # ### US$. This second format where the numbers are separated by spaces instead of commas and the dollar sign appears at the end of the string instead of the beginning with an extraneous US is incorrect. What SQL can I run to migrate all prices that are incorrectly formatted to the correct format?
Alternatively, this is a non-tech related one that GPT-4 gets correct that no other model does:
> In 388 BC, the boxer Eupolus of Thessaly defeated three opponents at the ancient Olympics. Funded by the men involved, several bronze statues of Zeus were erected, with inscriptions detailing the events. Why did the four men prefer not to have the statues around?
I like that one because Opus will tell you that you're wrong and mistaken, which is pretty funny while GPT-4 answer correctly with the actual context.
However, if I start with: "What is the earliest record of cheating in the olympic games?" Then all models get the question right. It's surprising that GPT-4 gets it right on the first go.
But it’s not by a ton, and it could be just less hyper-scale network infrastructure.
I don’t recall what Anthropic has raised but IIRC it wasn’t “peer with everyone” levels.
[1] https://www.bloomberg.com/news/articles/2023-12-21/anthropic...
[2] https://www.anthropic.com/news/anthropic-series-c
[3] https://www.aboutamazon.com/news/company-news/amazon-anthrop...