I've noticed the exact opposite, in my experience GPT3.5 is more likely to make code that doesn't work correctly, mess up URLs and other strings, or forget about certain variables and not include them. It's also more likely to make up non-existent libraries and other code.
GPT4 in my experience has been much more consistent with it's results. It also doesn't seem to lose the context as quickly. And it can deal with more niche libraries a lot better.
Due to the lack of consistently I've found it difficult to use 3.5
I haven't tried it with LangChain type stuff so maybe it's different there?
As a side note - I've been noticing their performance change a lot in between their "updates" and they mostly seem to be getting worse at following instructions :/