Not sure about this. Codex already uses structured diffs, which are much simpler conceptually and shouldn’t require full file rewrites. The GPT model is probably tuned for its apply_patch tool, too. I’d like to see some real world benchmarks/comparisons.