Serving a request to one of their APIs requires orders of magnitude more compute than your typical web service
Based on experience of BERT , yes maybe to get the best experience or to serve millions of users you need to run any model on compute intensive infrastructure , BUT if you just want to run for yourself and do some small testing you can very well download it from huggingface and elsewhere and run it on your laptop.
The amount of memory required to run these models is immense.
If you want a comparison, try running the largest version of the open source BLOOM model yourself.
"The Python code in this tutorial generates one token every 3 minutes on a computer with an i5 11gen processor, 16GB of RAM, and a Samsung 980 PRO NVME..."
[1] https://towardsdatascience.com/run-bloom-the-largest-open-ac...