It is worth noting that I had the perception, incorrect or not, that larger customers were being given priority to OpenAI APIs. It is good to see that you are reporting improvements.
Oh, they undoubtedly were. We have decent volume, but worked very hard in our prompting to minimize input and output tokens, so we don't have a big bill. I just think that in the past months they've added significantly more capacity and done a lot of work on their inference servers to make things run better.
Given Meta's obvious, and most welcome, open source play with LLaMA 2, it behooves OpenAI to be performant and accessible to everyone. For my application, I am finding that 70b LLaMA 2 models are very close to the inference quality of gpt-3.5-turbo.
It's gotten better for everyone in the last few months. It used to be a nightmare, but I haven't seen a timeout or rate limit error in a long time.