Fixing code is not the same as fixing a fundamentally broken architecture. The latter requires understanding that the architecture is broken in the first place and that understanding comes from experience.
And most of the time the architecture is not just broken, it's just meaningless because of so much accumulated error. If it was just broken then other comments would make sense, just throw a more capable model at it, but LLMs can't really reproduce shakespeare out of white noise just yet.