OpenAI has shown that these models at full power work great, so now they're trying to optimize for cost. I've gotten similar low accuracy responses from stuff it could handle a month ago.
It was kind of cringey when the model generated low accuracy nonsense the user detected that as "sentient." Come on