A $1200 GPU is all a reasonably capable engineer needs for each work stream. The cloud prices are a scam.
I run Qwen 3.8 27b on each of my six AMD AI PRO r9700 GPUs at 80tps decode each and they cranks for days with 256k context, doing complex kernel, compiler, debugging, enclave, bootstrapping, pentesting, hardening, and systems work full time.
~$1200-1500/ea on ebay.
Also ~40tps on my strix halo now but with room for several sessions at once.
Also using with Charmbracelet Crush with Froggerinc fixed chat templates.