Just to get OTP up for inference would require a very large spend. To use GPT-NeoX (20B parameters) for inference requires 45GB of vRAM minimally. Its hard for me to imagine using a 170B model for fun somewhere unless one has a large GPU farm or lots of money.
GPT-NeoX-20B was specifically targeted to fit on A40s, A6000s, and a pair of 3090 Tis. Anything larger than that is going to be a real struggle for people who don’t own computing clusters to use.