I have similar hardware - what specific version of the 3.8 Next model are you running and how many tokens/s are you seeing? 3.6 35B A3B gets about 68t/s for me so I've been sticking with that model for the mean time.
The thing with 3.8 next is that it uses a variant of ngrams. Part of the network is replaced by a lookup table you can store on a fast ssd.
In practice, you will be able to run models a bit bigger than 35B.
I have been using Q4 and running it with https://github.com/gufo-org/gufo which seems to be working pretty well. But I have to get off bazzite as my distro. I love the fedora atomic in principle but it's a pain for normal installs.