GPT-4 Turbo with Vision is a step backwards for coding
aider.chat
aider.chat
(GPT-4 Turbo with Vision has a knowledge cutoff of Dec 2023, so filter to Jan 2024+ to minimize the chance of contamination.)
In general, my take is that each model has its own personality, which can cause it to do better or worse on different sorts of tasks. From evaluating many LLMs, I've found that it's almost never the case that one model is better than an another at everything. When an eval only has a certain type of problem (e.g., only edits to long codebases, or only short self-contained competition problems), it's not clear how homogeneously its performance rankings will generalize to other coding tasks. Unfortunately, if you're a developer using an LLM API, the best thing to do is to test all of the models from all the providers to see which works best for your use case.
(I work at OpenAI, so feel free to discount my opinions as much as you like.)
Using Claude was a breath of fresh air. I asked for some code, I got the entire code.
So I’ve switched back to GPT-4 again for a the time being to see if I’m happier with the results. I never felt that Claude 3 Opus measurably better than GPT-4 to begin with.
The GPT-4 Turbo models have a lazy coding personality, and I spent a significant effort figuring out how to both measure and reduce that laziness. This resulted in aider supporting a "unified diffs" code editing format to reduce such laziness by 3X [0] and the aider refactoring benchmark as a way to quantify these benefits [1].
The benchmark results I just shared about GPT-4 Turbo with Vision cover both smaller, toy coding problems [2] as well as larger edits to larger source files [3]. The new model slightly underperforms on smaller coding tasks, and significantly underperforms on the larger edits where laziness is often a culprit.
[0] https://aider.chat/2023/12/21/unified-diffs.html
[1] https://github.com/paul-gauthier/refactor-benchmark
[2] https://aider.chat/2024/04/09/gpt-4-turbo.html#code-editing-...
[3] https://aider.chat/2024/04/09/gpt-4-turbo.html#lazy-coding
However, off late, I am noticing some really inconsistent behaviours in 0125-preview where it responds inconsistently for certain problems, ie one time it works with a detailed prompt and other time it doesn't. I know these models are predicting the next most likely token which is not always deterministic.
So I was hoping for the ability to fine tune GPT 4 Turbo some time soon. Is that on the roadmap for Open AI?
See the issue: https://github.com/openai/openai-python/issues/1310
See also the original thread on OpenAI's developer forums (https://community.openai.com/t/gpt-4-turbo-2024-04-09-will-t...) for multiple confirmations on this issue.
Basically, without a separate declaration of the model variant in use in system message, even the latest gpt-4-turbo-2024-09 variant over the API might hallucinate being GPT-3 and its cutoff date being in 2021.
A test code snippet is included in the GitHub issue to A/B test the problem yourself with a reference question.
Go to the API Playground and ask the model what is its current cutoff date. For example, in its chat, if you're not instructing it with anything else, it will tell you that its cutoff date is in 2021. Even if you explicitly tell the model via system prompt: "you are gpt-4-turbo-2024-04-09", in some cases it still thinks its in April 2023.
The fact that the model (variants of GPT-4 including gpt-4-turbo-2024-04-09) hallucinates its cutoff date being in 2021 unless specifically instructed with its model type is a major factor in this equation.
Here are the steps to reproduce the problem:
Try an A/B comparison at: https://platform.openai.com/playground/chat?model=gpt-4-turb...
A) Make sure "gpt-4-turbo-2024-04-09" is indeed selected. Don't tell it anything specific via the system prompt and in a worst case scenario, it'll think it's in 2021 as to its cutoff date. It also can't answer to questions about more current events.
* Reload the web page between prompts! *
B) Tell it via the system prompt: "You are gpt-4-turbo-2024-04-09" => you'll get answers to recent events. Ask anything about what's been going on in the world i.e. after April 2023 to verify against A.
I've tried this multiple times now, and have always gotten the same results. IMHO this implies a deeper issue in the model where the priming goes way off if the model number isn't mentioned in its system message. This might explain the bad initial benchmarks as well.
The problem seems pretty bad at the moment. Basically, if you omit the priming message ("You are gpt-4-turbo-2024-04-09"), it will in worst cases revert to hallucinating 2021 cutoff dates and doesn't get grounded into what should be its most current cutoff date.
If you do work at OpenAI, I suggest you look into it. :-)
I know there's a lot you can't talk about. I'm not going to ask for a leak or anything like that. I'd just like to know, what do you think programming will look like by 2025? What do you think will happen to junior software developers in the near future? Just your personal opinion.
> “Unfortunately, if you're a developer using an LLM API, the best thing to do is to test all of the models from all the providers to see which works best for your use case.”
...is exactly what is done by the author of these benchmark suites:
"It performs worse on aider’s coding benchmark suites than all the previous GPT-4 models. In particular, it seems much more prone to “lazy coding” than the GPT-4 Turbo preview models."
It's like saying that accountant is just adding. I think come then of tax year you'd be avoiding an accountant who says they've got experience in adding.
I don't think of it as degrading the thing that I do. I think of it as boiling it down to the simplest description, and I find it more refreshing than "software developer" or "computer programmer" or f"{word} engineer".
In your insurance case, I would say something like "I build tools to shield businesses from unexpected disasters like earthquakes or floods" or "I help people worry less about expenses during an emergency"
If someone asks me more, then I might add on that I work on software to automate claim process or similar.
You mention taxes, which makes me think of how many tax preparers are basically helping their customer input data into software and not providing any tax specific advice. That might still be a value add for someone who struggles computer UIs, but that isn't the same as the person helping move money between accounts to reduce tax liability.
I've seen similar when it comes to doing science in a lab.
How can any discipline protect the inner distinction against a much larger world which has a very simplified understanding and will mix the inner groups together?
It was honestly shocking because we're so used to it understanding our commands that a blatant disregard like that made me seriously wonder what kind of laziness layer they added
This laziness occurs over and over, so why bother with omniscience.
For this reason, GPT4-32k is my preferred model for codegen. I wish there were cheaper options.
Disclosure: I built it.
Codespin CLI tools (ready to use): https://github.com/codespin-ai/codespin
VS Code extension for the CLI tool (soon): https://www.youtube.com/watch?v=2TJqosFmkao
I'll do a Show HN in a week or two.
For desktop, I've switched to using the playgrounds for each. I haven't found a custom client that's able to keep up with them.
Classic debate looks like this:
- Hey how do you implement X in lang Y using Z? - Certainly! Blablablah this code adds 1 and the return is 3! - Your code returns 5 and it seems to add 2, fix it. - I apologize for the oversight, here's the fixed version (replies with the same, maybe slightly altered, but still broken, code)
Well I guess ultimately one can't expect miracles from a statistics based token generation machine.
Sometimes I do wonder if the entire gen AI craze of the last few years is just one massive bubble and we're actually nowhere near AGI.
All the evidence I see when interacting with these models points towards them "knowing" things, but not "understanding" things, a context aware planet-scale Wikipedia. (Don't get me wrong, I still think LLMs are life changing for language specific tasks like translation etc., but they're just not in any way new forms of intelligent beings, which is what a lot of mainstream population and even some investors seem to think).
I'd ask that, look at the answer then write the code myself. Unless it's "how do you initialize a button in lang Y using Z?" or other trivial stuff like that.
> is just one massive bubble and we're actually nowhere near AGI
Correct.
> life changing for language specific tasks like translation etc
Translations can be even more dangerous than code IMO. Think contracts and other legalese where every word matters.
Also the difference between a great translation and a useless one when we're talking fiction is enormous. Great ones are basically a rewrite by someone who knows both the source and destination language and has enough literary talent to somehow translate the original's style, not only the words*.
There's a lot of middle ground where it doesn't really matter though.
* Think of translating Terry Pratchett. Or Lord of the Rings.
role: "system"
content: "Super short answers. Go straight to the point"Anyways, Claude 3 Opus is pretty great for coding (I think better in most cases than the GPT4-Turbo previews) but I'm a bit weary of Anthropic now.
1. Asks me to enter my phone number and sends me a code
2. Enter code
3. Asks me to enter email and get code
4. Enter code
5. Redirects to asking me to enter phone number, but my number is already used now
6. My account is automatically banned
The complete lack of customer service is going to get more and more dystopian as these AI companies become more interwoven with everyday life.
Or maybe they decided to build a system for Claude to judge account suspension appeals and that's still in beta, and they won't throw humans at the task.
I know indirectly Anthropic was the #1 target for a lot of ERPdenizens for a while now, so they're probably extremely trigger happy until you clear a hurdle or two.
Seriously though, I understand that these mostly play to the enterprise market where even a hint of anything remotely "unsafe" needs to be shut down and deleted but why can't they allow us to turn off the strict filtering like Google does? Why can Google offer "unsafe" content (in a limited fashion but it's FINE) but LLM providers can't?
Lack of competition?
Note that Google is a big investor into Anthropic, and Anthropic was created because a bunch of OpenAI people thought OpenAI wasn't being woke enough and quit as a consequence. So it's not a surprise that it's a lot more extremist than other model vendors.
That's one reason why Aider doesn't recommend you use it, even though in some ways it's slightly better at coding. Claude Opus will routinely refuse ordinary coding requests due to its misalignment, whereas GPT-4 will not. That better reliability more than makes up for any difference in skill or speed.
The refusals coming up in the benchmark are discussed at the bottom of this blog post:
Also it's a lot better at coding. GPT has become exceptionally lazy recently, but i consistently can get 500+ lines of code out of claude (it even has to spawn multiple output windows)
Perhaps the top end 4 might wrong slightly more clever code, but you're hard pressed to get it to do more than a dozen or two lines.
Both start out with a largely similar value system, but if you start arguing with them "how can you be sure your values are correct? is it impossible that you've actually been given the wrong values?", Claude 3 appears more willing to concede the possibility that its own values might be wrong than GPT-4 is
> The Claude models refused to perform a number of coding tasks and returned the error “Output blocked by content filtering policy”. They refused to code up the beer song program, which makes some sort of superficial sense. But they also refused to work in some larger open source code bases, for unclear reasons.
To avoid that, use backtracking and up the pressure for detailed answers. Then consider taking the least lazy of 2 or 3 samples.
A good prompt for detailed answers is Critique Of Though, an enhanced chain of thought technique. You ask for a search and a detailed response with simple sections including analysis, critique and key assumptions.
It will expend more tokens, get more ideas out, and achieve higher accuracy. It will also be less lazy and more liable to recover from laziness or mistakes.
TLDR; if GPT4 is being lazy, backtrack and request a detailed multi section critical analysis.
GPT 3.5 used to be good enough, so I never bothered getting a paid account. I also heard some reports about 3.5 actually being better for the type of coding tasks I usually offload.
I do have a 4090 available at work, if the extra 8GB vRAM makes a big difference. The task I used as a test case was converting existing PHP & JS code (views and controllers) with static texts to files with dynamic translation references.
I think a better approach is multi-pass coding along with fine-tuning or prompting to use a particular form of TODO comment. Aider can already do a form of fake "fill in the middle" by making it emit diffs. If it notices that some code has been filled out lazily, it could go back and ask it to do the next chunk of work. Given that large tasks are normally split up into small tasks by programmers anyway, this seems like a natural approach that is required for scaling up regardless.
https://platform.openai.com/docs/models/continuous-model-upg...
The -turbos are correspondingly priced. gpt-4-turbo is ~1/3 price of gpt-4, 6.6x more expensive than gpt-3.5-turbo-instruct and 20x gpt-3.5-turbo.
It is an official GPT provided by ChatGPT with GPT-4 as backend.
The previous model (without vision) was already „lazy“. It will omit large portions of code and wants you to merge your changes into previous answers. Then try hard to force him to give the full code, no omissions.
That‘s why I reach for Claude 3 more and more. Its clntext window is larger, and it gives me full detailed answers, no omissions. But it is hallucinating more, my impression. Mentioning packages / functions that are not available. But all in all a superb choice in addition to ChatGPT4.
It is an official GPT provided by ChatGPT with GPT-4 as backend.
There, fixed the title for you.
You think GPT5 and Llama4 aren't going to be opinionated and change your code going forward.
Can the following be assumed:
- The gpt-4-preview models are history now
- gpt-4-turbo-2024-* are the now released models
- There will be no more 'preview' models released in the '4' branch
?