A paper from a week ago found that models trained on multiple data modes perform an order of magnitude better than text-only models of the same or even larger size.
Being able to perform more advanced types of zero shot learning tasks would be comparable and further the accuracy on those tasks can be evaluated
With 16k and some other techniques, I’m guessing it could write a custom CMS database backed web application.
more literally and correctly, it’s the maximum number of tokens in the input and output, combined, where a token is 4/3 of a word
So we’re shifting from 5K words maximum to 40K (per sibling comment, who pointed out 32K context leaked as well)