4o added image analysis.
The o-series starting at o1 improves on 4o as per the margins in these charts: https://openai.com/index/learning-to-reason-with-llms/
I'll have to wait and see about o3, because only the mini model is out yet.
[0] https://law.stanford.edu/2023/04/19/gpt-4-passes-the-bar-exa...
[1] https://ai.nejm.org/doi/full/10.1056/AIdbp2300192
[2] https://www.betonit.ai/p/gpt-4-takes-a-new-midterm-and-gets
I get that feeling too (in both directions) but this vague and hard to quantify sensation is not what I'd suggest in response to your clearly stated question:
> And what has enormously improved since ChatGPTs launch?
Which is, I think, answered by the things I listed.
But you insist on being obstinate. ChatGPT advised me to disengage from this conversation.
ChatGPT did not ace the bar exam -- it was basically percentile graded against a group of people who mostly failed. If compared to real lawyers, it was 15th percentile on the essay portion
[0] https://law-ai.org/re-evaluating-gpt-4s-bar-exam-performance...
15th percentile of passes, on the weakest aspect, is still a big improvement over "not passing". That improvement is what I wish to highlight.
(The observation that 48th percentile (lowest overall from your link, let alone 15th for essays) of passes corresponds to 90th percentile of all exam takers, suggests that perhaps too many humans are taking the exams before they're ready).
I’m sure if you need to make a to-do list in react it’s like magic (until the app gets complicated). In real world use, not so much.
(Also I have often code reviewed PRs from people who are heavy users and surprise surprise - their output is trash and very prone to bugs or being out of spec.)
- reverse engineering: when fed assembly (or decomp or mock impl), it's been consistently been able to figure out what the function actually does/why it's there from a high-level perspective. Whereas ChatGPT merely states the obvious
- very technical C++ questions: DSR1 gives much more detailed answers, with bullet points and examples. Much better writing style. Slightly prone to hallucinations, but not that much
- any controversial topic: ChatGPT models are trained to avoid these because of its "safety" training
ChatGPT is a bit better (and faster) at writing simple code and doing some math faster, but that's it.
(obviously, common sense about what to share and not to share with these chatbots still apply, etc.)
There's lots of fiddling with these models. I found Claude 3.5 Sonnet to be superior to both GPT-4o and o1-preview in around 99% of the things I do; I only started comparing it against o3-mini, and right now it's a mixed bag. Then again, I tend to develop and refine specific prompts for Sonnet, which I haven't for o1-preview and o3-mini, so that could be a factor. Etc.
Yes, well, I live in the EU and thus can avoid US work hours and Chinese peak hours. I think availability has been a bit better since they disabled websearch (also I noticed DSR1 half a week before it made the mainstream news).
> There's lots of fiddling with these models.
Agreed
--
The other is simply using the technology for what it's good for, observing that it's slowly, incrementally improving at tasks that it was already capable of since the major breakthrough, and acknowledging its limitations.
Incremental improvements don't give us any assurance that another major breakthrough is waiting around the corner.