But it makes little to no sense as long as API prices are what they are. Except for maybe privacy reasons.
But it makes little to no sense as long as API prices are what they are. Except for maybe privacy reasons.
Going local has as opened up a world of use-cases I never would have entertained the idea of on metered/cloud usage. Privacy is a large part of it but, I also no longer think twice about whether to send a prompt or not based on the psychology of it costing money.
Cached input tokens on local inference are free, so I don’t care about running sessions up to 500k tokens and hundreds of turns (it’s rarely useful, but DSv4 remains surprisingly coherent up there)
owning a few GPUs is a lot cheaper than supercars.
I don't use it for local inference so much. I use it to learn.
I also use it as my daily driving Aarch64 development system.
Aside it's also very cool what else can be done with unified GPU memory, once you realize you have it...