We’ve noticed a similar quality to GPT-4o, with the same level of prompting needed to get the same output for harder tasks.
As a simple example from yesterday, we were generating a GitHub Action using GPT-4o and running into an issue.
The GitHub action does a diff of your code against `main` and then passes that diff to an LLM for a simple vulnerability analysis, returning a severity of high|medium|low plus a description.
Both GPT-4o and Claude 3.5 Sonnet kept getting caught up in the same problem, which is that the workflow they generated would log the contents of PR, which itself contained said workflow as YAML. This then caused the YAML to parse incorrectly.
A small example, but the issue was obvious to a developer reading the output. Both LLMs struggled to “understand” the issue, even when they were directly told what to do (they kept logging the diff for debug, which would continue to break it).
We got to a solution with both LLMs, but it required us to be quite explicit in our prompting.
Caveat: both GPT-4o and Claude 3.5 Sonnet were accessed via a self-hosted interface that hits the API directly, so your experience through the UIs may be different.