FWIW, we've been live with GPT-3.5 turbo since May and it's improved a LOT.
Latency is down consistently across the board and we haven't seen a single "429 model overloaded" error in the past month: https://twitter.com/_cartermp/status/1686894576202907651/