We sure did. It's a great writer, better in a harness, will process lots large context, but complex reasoning with convoluted rules and low hallucination tolerance? That's still larger model territory.
There's still 3.1 Pro though, as ancient as it sounds now
Yup. The problem is that it's bizarrely still not in General Availability status.
sounds like you might need to beef up your harness first, and run multi-agent verification loops
That's fine if you're doing interactive dev tasks, but we're in the large volume, cost effective, big inputs, business still with low error tolerance business, and tuned the heck out of what we can get with minimal fix cycles. Millions of cases at hundreds of thousands tokens each - after all the prefiltering by cheaper models - and the tasks still need them to do convoluted reasoning. So 'usually get it right the first time' is a big part of the cost equation.