Calling this garbage is absolutely wild. The authors make it very clear that this is optimized for throughput and not latency. Throughput focused scenarios absolutely do exist, editorializing this as "running large language models like ChatGPT" and focusing on chatbot applications is the fault of HN.
It's also a neat result that fp4 quantization doesn't cause much issue even at 175b, though that kinda was to be expected.