We tested it across 100 unsaturated coding and engineering environments. Both Astra and 6.1-Sol are pretty comfortably ahead of Opus 5.5 in these types of evaluations, and both end up being cheaper than Opus via API usage. 6.1-Sol is also cheaper than Sonnet 5.5 and much smarter. The only verifiable domain Anthropic seems to be clearly ahead is chemistry (and perhaps also some unverifiable domains like being pleasant to work with, since GPT-6 models have a tendency to be low initiative beyond what the prompt tells them). Highly recommend using the OpenAI Flex endpoint for any API work.
Data at https://gertlabs.com/rankings