Especially because these models <think> for a long bit before actually answering, so they generate for a longer period. (to be clear - I find the <think> section useful, but it also means waiting for more tokens)
Personally - I end up moving down to lower quality models (quants/less params) until I hit about 15 tokens/second. For chat, that seems to be the magical spot where I stop caring and it's "fast enough" to keep me engaged.
For inline code helpers (ex - copilot) you really need to be up near 30 tokens/second to make it feel fast enough to be helpful.