The model being tested is 18k as configured.
I didn't expect this to make the 5090 to look like a good deal.
I didn't expect this to make the 5090 to look like a good deal.
It'd be silly to buy the 18k model to run a tiny model like Qwen 27B. You use models like GLM Flash and Qwen Next which won't fit on a single 5090.
(Each task needs its own context, but the (e.g.) 27B of constant parameters isn't duplicated).