40-80 tok/s is unusable to you? Ok.
If you like having a box with 8-12 fans blasting hot air and noise into your office all day, nobody's stopping you.
If you like having a box with 8-12 fans blasting hot air and noise into your office all day, nobody's stopping you.
the 40-80 tok/sec is only for initial prompt processing, and with the "medium" models, like Qwen3.6:27b. The actual token generation is in the 10 token/second Thats very slow. And your Macbook pro will stop being a LAP-top, because it will get very warm.
Meanwhile, my 2x3090s happily crank out ~100 tok/sec generation. Oh and I can run 100 tok/sec on my phone as well, because I can just access ollama on my home desktop over ssh from termux.