A mid-level human engineer can iteratively fix the mistakes they start making, instead of that just being the final result. Can GPT-4 do that?
Hooking it up to automatically run the code in question and examine the output is a trivial undertaking - many folks have already done this.