I think the big reason why BERT and T5 have fallen out of favor is the lack of zero shot (or few shot) ability.
When you have hundreds or thousands of examples, BERT works great. But that is very restricting.
When you have hundreds or thousands of examples, BERT works great. But that is very restricting.
Is this because the final represention in bert style models more globally focused, rather than being optimized for next token prediction?