I don't really understand what the value in posting these kinds of takeaways about using GPT3.5 here is. GPT4 is significantly better, and improved models are coming. There's just not a lot of point to benchmarking 3.5 when likely every issue you've pointed out is solved by 4.