Are they really publishing a paper based on GPT3.5 in July 2024? I am not sure these results are relevant in any way today.
Edit: Just for reference. The best model for coding today (according to most benchmarks) is Claude-3.5-Sonnet which is freely accessible. Also GPT-4o is freely accessible and is still vastly better than GPT-3.5.
The lm sys arena coding leaderboard (https://chat.lmsys.org/?leaderboard) lists sonnet-3.5 and gpt-4o jointly on #1 and GPT-3.5-Turbo on #35. You can freely download and run LLMs locally on your machine that are significantly better than GPT-3.5, for example Mistral Codestral.
There is really no reason to accept any results on GPT3.5 for relevant today. This is as if you were complaining that a computer from the 00ies is not running <recent operating system> well.