Every model with ~4B parameters runs perfectly fine on even a Geforce 1070 Mobile GPU with 8Gb of memory.
If you have some patience you can probably go a little crazy and run a model with ~27B parameters on a Radeon 890M with 32Gb of memory as well (means you'll probably have to get about 96Gb of system memory if you want to get some work done too, but oh well).
In theory you could even run a model which fits in 64Gb of video memory on that "little" GPU (with 128Gb of system memory).
No, you can't run something like Grok 2 (which has quantified models starting with 82Gb in size and going up) but why on earth would you ever want to run something like that locally?