Even during that lull between GPT3 and DALL-E/CLIP, there was tons of truly wonderful advances in AI...
In the important way that the AI winter originally referred to though, no, there doesn't seem to have been any progress towards AGI.
I do think the last few years have been more productive than previous periods in advancing narrow AI, and to be fair to those researchers who just get on with the work, it is not on them if the advances are over-sold by others.
I bet we're closer than most people think. Instruct GPT-3 can do semantic tasks just as efficiently as DALL-E 2 can draw. NLP tasks that took whole teams multiple years can be simply described in a few words and they work right away.
The entry barrier to implement new tasks will get very low. The large models will be the new operating system. This means more investments and data, leading to new improvements.
I believe GPT-3 is already close to median human level on most semantic tasks that fit in a 4000 token window. I'm researching how to use it right now for a variety of tasks, it just works from plain text requirements with no training.
But there's a quantum leap or two from (say) mindlessly producing comments that can occasionally fool readers on HN (as it has been used to do in the past) to it consciously joining in the conversation of its own volition and curiosity, then zoning out on Netflix while half-worrying about the future for GPT-4 jr and idly planning it's next server room refit.
I'm actually more curious if we could parse the underlying logic that ultimately it emulates to merge those images together.
It 'looks like' something kind of sophisticated is being modelled with AI but there's some nice algorithms hidden in there.
The training principle of CLIP is very simple, but intuitively understanding how the diffusion prior maps between semantically similar textual and visual representations is a bit more unclear (if that's even a well-formulated question!)
More like 'averaging them' and finding variations from vast inputs.
Which is more a long the lines of what I mean.
> While our model can render a wide variety of text prompts zero-shot, it can can have difficulty producing realistic im ages for complex prompts. Therefore, we provide our model with editing capabilities in addition to zero-shot generation, which allows humans to iteratively improve model samples until they match more complex prompts. Specifically, we fine-tune our model to perform image inpainting, finding that it is capable of making realistic edits to existing images using natural language prompts.
Unless that only applies to GLIDE and not to DAL-E?