You can also use int8 quantization which comes for free with huggingface and isn’t a huge cost to accuracy. Which cuts the memory usage almost in half again. (almost in half because some layers might stay 32 bit if they make a tiny part of the model)