Just to get OTP up for inference would require a very large spend. To use GPT-NeoX (20B parameters) for inference requires 45GB of vRAM minimally. Its hard for me to imagine using a 170B model for fun somewhere unless one has a large GPU farm or lots of money.