These performance changes seem pretty inevitable when OpenAI is going to continually update the models. Short of versioning every iteration of the model I don't see how developers can avoid these issues. The solution seems to be to implement better telemetry where these APIs are used in production. I've been working on a tool to help with this - www.getcontext.ai