ParentFull threadwsgeorge·I would guess that a) they assume their users will have a lot of GPU ram, b) actual running costs will depend on what inference engine/framework you're using (for example, some GGUF quants are very cheap).View on HN