TL;DR?
(Yes, GPT-4 is prob better... but by how much and on what?) A table would've been easier
(Yes, GPT-4 is prob better... but by how much and on what?) A table would've been easier
The conclusion is that neither 3.5 nor 4 are good enough because for anything none trivial they generate code that is often subtly wrong. Might still speed up somebody new to the language/project/learning or I would say: with additional tooling/plugins/"prompt engineering"/tinkering the author might get useful results.
from https://github.com/E-xyza/Exonerate/blob/master/bench/report...
(I believe the author is significantly underestimating the pace of progress)
Specific numbers are at https://github.com/E-xyza/Exonerate/blob/master/bench/report.... GPT-4 does significantly better.