A 4-bit quantized version will require (still a whopping) ~100GB of RAM using tools like Ollama and Llama.cpp
If you want to try it with Ollama on macOS (keeping in mind you'll need the new Mac Studio with 192GB of memory) this command will work:
ollama run falcon:180b
Right now it gets about ~5 t/s.. there's quite a bit of work being done to improve inference speeds like speculative sampling which uses a smaller model for a subset of tokens