Interesting that they are broadening instead of deepening the text side of GPT - maybe they are running into problems on that side and are building the feature set with low hanging fruit on adjacent areas.
Interesting that they are broadening instead of deepening the text side of GPT - maybe they are running into problems on that side and are building the feature set with low hanging fruit on adjacent areas.
ChatGPT struggles with intuition about the world, if it has image/video knowledge than it can potentially reason about what happens when you flip a plate upside down.
There is an interesting “takeoff” point where models are able to generate enough “interesting” data to train themselves. Asking chatGPT for ml training data on various tasks makes me think we are close to that inflection point.
Dalle/GPT: that's not right, there should be 5 fingers on this clock.
It turns out that the image models already understand the low-level details of images well, and it's high-level reasoning/composition that is screwing their samples up - which is due to the poor text embedding they receive as a summary. (You can't generate a good image of 'a blue ball left of a red ball' if your text embedding is confused what 'left' is, no matter how high quality you are able to generate 'blue ball' or 'red ball' individually.) So, I would not be surprised if hands are improved by simply using a much better text model. (And why not plug in GPT-3 or PaLM for even better results...)
That said, hands are also like text-inside-images in being small, highly variable, and intrinsically difficult. And we know that the larger models solve text rendering simply by scaling up an OOM or 2 past SD. (You can see this from the text samples in the papers, and also more recently from Stability's Deep Floyd samples they've been teasing on Twitter.) So either way, scaling will probably solve hands soon.
One of their new models, PaLM-E, does show some positive transfer, so it's definitely a potential path for increased capabilities.
Also, amazingly, transfer learning works better than anyone could have hoped! As an analogy, imagine that you could show that football players learned musical instruments 50% quicker than non-athletes. That would be amazing!
But LLMs can write music and training them on music makes them better writers. It's almost like intelligence & skills are completely general!
I know the entire AI industry on the wrong track, but there's nothing that can be done about it. Sometimes you just have to wait for the crash and burn, and then things can change. It's like being at the ocean and trying to stop a wave from crashing after it has already formed, you just have to let it resolve itself of its own volition. As Einstein said, stupidity is infinite.
Being able to perform more advanced types of zero shot learning tasks would be comparable and further the accuracy on those tasks can be evaluated
With 16k and some other techniques, I’m guessing it could write a custom CMS database backed web application.
more literally and correctly, it’s the maximum number of tokens in the input and output, combined, where a token is 4/3 of a word
So we’re shifting from 5K words maximum to 40K (per sibling comment, who pointed out 32K context leaked as well)