I own an 8 GPU cluster that I built for super cheap < $4,000. 180gb vram, 7 24gb + 1 24gb. There are tons of models that I run that's not hosted by any provider. The only way to run it is to host myself. Furthermore, the author has 39 tokens in 6 seconds. For llama3-8b, I get almost 80 tk/s and if parallel, can easily get up to 800 tk/s. Most users at home infer only one at a time because they are doing chat or role play. If you are doing more serious work, you will most likely have multiple inference running at once. When working with smaller models, it's not unusual to have 4-5 models loaded at once with multiple inference going. I have about 2tb of models downloaded, I don't have to shuffle data back and forth to the cloud, etc. To each their own, the author's argument is made today by many on why you should host in the cloud. Yet if you are not flush with cash and a little creative, it's far cheaper to run your own server than in the cloud.
To run llama-3 8b. A new $300 3060 12gb will do, it will load fine in Q8 gguf. If you must load in fp16 and cash is a problem a $160 P40 will do. If performance is desired a used 3090 for ~$650 will do.