Releasing model weights has been common for a long time, it was OpenAI that regressed and refused to release its GPT models until it released GPT-3 hidden behind a semi-public API. The google BERT models are released pre-trained in full (as mentioned in the article), you can easily play around with them on your own.
BERT isn't a forward LM like GPT, so it's a little less easy to play with for the uninitiated - although there are papers showing text generation with BERT.
At least that's my understanding. Feel free to correct me if I'm wrong.
GPT-2 wasn't really feasible for most actors to train from scratch, otherwise it would have been released by a third party. GPT-3 is technically feasible for a single actor to do inference with, although not really really.
Still in the realm of doing inference in homelab territory (barely).
Where's this heuristic from? Seems handy if true.
Edit: It probably is, considering you only need weights to be fp16, while intermediate layers can be fp32 and reuse these memory.