Code editing benchmarks for OpenAI's "1106" models
aider.chat
aider.chat
Is the new, huge context size going to change how internals work? Seems like if you were being careful and efficient before, maybe that can be traded for quality somehow.
Aider is helping me write rust, which I'm very new to. It's doing a reasonable job, despite people saying GPT4 does a bad job at rust. It's a very small project--LLM based, of course! It makes frequent errors, but it can fix them by just feeding it `/run cargo check`--sometimes a few iterations.
It's a cool workflow, and I'm already a terminal nerd, having been using a tmux/cloud dev environment for years. Aider's ergonomics are perfect for me. I was previously making use of chatblade in the terminal to pipe files into etc.
I'm looking forward to see how good aider gets "for free" as the models get better.
Anyway, I'll be using it, and watching for any news you might have. Thanks again!
The new large context size certainly opens up some possibilities. Most LLMs start to get distracted/confused if you just put everything and the kitchen sink into their context window. I expect that would be a concern if you just dump in 128k tokens worth of code, most of which isn't relevant to the task at hand.
I imagine we still want to focus GPT on the relevant code and use the larger context window to provide more "code context" along the lines of aider's repository map. Right now the repo-map is optimized to fit within 1k tokens by default. It seems likely that the new GPT-4 Turbo model should be given a much larger map than that.
I'll be experimenting and optimizing aider for this new model over the coming days and weeks.
These results are preliminiary. OpenAI is enforcing very low rate limits on the new GPT-4 model. I will update the results on this page as quickly my rate limit allows.
Either way, the new model is looking very promising.
GPT-3.5 is only able to reliably edit a file by returning a whole new copy of the file with the edits included. This is the "whole" edit format.
GPT-4 is able to use a more efficient "diff" edit format, where it species blocks of code to search and replace.
All of this is described and quantified in more detail in the original aider benchmarking article.
Also, I'd not realised 3.5 would be faster too, that's huge.