At some point the models will produce code with a lower error rate than existing libraries.
Review and testing.
Reviewing is easier when there is less code (i.e. libraries are in use)
These were tricky problems that were small scope - I've picked them so I could easily provide it to GPT for review.
So I doubt larger context window will do much.