The first benchmark chart of LLama.cpp is effectively a memory bandwidth chart, unsurprisingly. Each of the cards except the 5090 gets almost exactly 0.1 token/s per GB/s memory bandwidth.
Mem GB/s tok/s tok/s pr GB/s
RTX 5090 1792 158.77 0.089
RTX 4090 1008 100.51 0.100
RTX 3090 936 90.88 0.097
RTX 4080S 736 77.08 0.105
RTX 4080 717 74.56 0.104
RTX 4070TiS 672 70.83 0.105
RTX 4070S 504 54.59 0.108
RTX 4070 504 54.47 0.108
The 5090 is an outlier in that it has 78% more memory bandwidth than the 4090, and "only" gets 0.09 token/s per GB/s. So it seems to start running into compute limitations.