I understand the spirit in this line of criticism, but I think it's easy to muddle the timelines and feel as if things "aren't moving," when in fact, the pace of research and improvement is great.
For context:
- GPT 2 was released in Feb 2019
- GPT 3 came out roughly 18 months later in 2020. It was a huge jump, but still not "usable" for many things.
- InstructGPT came out roughly 18 months later in early 2022, and was a huge advancement. This is RLHF's big moment.
- About 10 months later, ChatGPT is released at the end of 2022 as a "sibling" to InstructGPT. It's an "open research preview" at this point. This is around the time OpenAI starts referring to certain models as being in the "3.5 family"
- GPT-4 comes out in March 2023, so barely 2 years ago now. Huge jumps in performance, context window size, and it supports images. This is around the time ChatGPT hits 100 million users and is really becoming a reliable, widely adopted tool. This is also the same time that tools like Cursor are hitting the market, though they haven't exploded yet. Models are just now getting "good enough" for these kinds of applications
- GPT-4-Turbo comes out in November 2023, with way larger context windows and lower pricing.
- About 12 months ago, GPT-4o released, showing slightly increased performance on existing benchmarks over 4, but now with state-of-the-art audio capability support for something like 50 languages.
- 5 months ago, o1 releases. This is a big moment for scaling compute at test time, which is a major current research direction in ML. Shows huge improvements (something like 8x over 4) on some math/reasoning benchmarks. Within months, we have o3 and o4, which substantially improve these scores even further.
- In February of this year, we get 4.5, and then months later, the confusingly named 4.1, which shows improvements over 4o.
So to be clear, in 2019 we had an interesting research project that only a few people could tinker with.
18 months later, we had a better model that you could play with via an API, but was still a toy.
It takes more than two years to go from that to ChatGPT, and a few more months (nearly 3 years total) to get to the "useful" version of ChatGPT that really sets the world on fire. It took roughly 4 + 1/2 years to go from "novelty text generation" to "useful text generation".
In the 2 years since then, we've gotten multimodal models, a new class of reasoning models, baseline improvement across performance, and more. If anything, there is more fundamental research and wider variety of directions now (the kind of stuff that shifts paradigms) than before.