How does the speed of this model compare to other LLMs? I see lots of accuracy benchmarks, like HellaSwag, but are there performance benchmarks out there as well?
It entirely depends on the speed of your hardware, but roughly we'd expect it to be 3.5 times slower than Falcon 40B.
Either on a standardized set of hardware or relative to other models. Performance benchmarks exist for all sorts of compute intensive things, so surely there’s at least one for LLMs?