Assuming you used single precision the model is 350 gigabytes (175 billion * 2 bytes). For fast inference the model needs to be in GPU memory. Most GPUs have 16GB of memory, so you would need 22 GPUs just to hold the model in memory and that doesn't even include the memory for activations.
If you wanted to do fine tuning, you would need 3x as much memory for gradients and momentum.