The memory constraint is certainly an issue, but I think they can be overcome with some software hacks (for a speed penalty, of course.) For example, gradient checkpointing and gradient accumulation might help:
https://medium.com/tensorflow/fitting-larger-networks-into-m... https://medium.com/huggingface/training-larger-batches-pract...