Lazy coding benchmark for GPT-4-0125-preview
aider.chat
aider.chat
Today, we are releasing an updated GPT-4 Turbo preview model, gpt-4-0125-preview. This model completes tasks like code generation more thoroughly than the previous preview model and is intended to reduce cases of “laziness” where the model doesn’t complete a task.
With that in mind, I am benchmarking the new model using aider’s existing lazy coding benchmark. I've shared some preliminary results.
Overall, the new gpt-4-0125-preview model does worse on the lazy coding benchmark as compared to the November gpt-4-1106-preview model:
1. It performs much worse when using the unified diffs code editing format.
2. Using aider’s older SEARCH/REPLACE block editing format, the new January model outperforms the older November model. But it still performs worse than both models using unified diffs.