LLM Falcon 180B Needs 720GB RAM to Run
kaitchup.substack.com
kaitchup.substack.com
Still, 180B is a best to load up.
Zero copy does not play a role here - the problem is how the model is being loaded in the steps that are shown in the article and they have nothing to do with Safe tensors' zero copy ability. Check out the article by Hugging face to see how to do it with only 360GB instead of 720:
https://huggingface.co/docs/accelerate/usage_guides/big_mode...
Best -> Beast
The title from the article I see it now is "Falcon 180B: Can It Run on Your Computer?"
The content does mention 720GB RAM, but quickly clarifies that it can be reduced:
QUOTE:
Step 1 and step 2 are the ones that consume memory. In total, you would need 720 GB of memory available. This can be CPU RAM, but for fast inference, you may want to use GPUs, e.g., 9 A100 with 80 GB of VRAM. This wouldn’t be enough for batch inference.
With either CPU RAM or VRAM, that’s a lot of memory. Fortunately, these requirements can be easily reduced.
However, it transpires Falcon 180B uses a proprietary license “based upon” the Apache license. Despite Falcon's claims to the contrary, it is not open source.
And that's disappointing because Falcon 40B is actually open source (and Apache licensed).
Edit: plus, the industry acts like models aren't derivative works of the training data. If true the source license is less important than the model!
The Falcon 180B license is special, even in the context of "open" LLM models:
https://huggingface.co/spaces/tiiuae/falcon-180b-license/blo...
> 5.3. The Acceptable Use Policy may be updated from time to time. You should monitor the web address at which the Acceptable Use Policy is hosted to ensure that your use of the Work or any Derivative Work complies with the updated Acceptable Use Policy.
This is basically "don't use it" for any business that retains a lawyer.
Feels very EEE to me.
This isn't as big of a deal as people try to claim it is.
LLama 2 can be used for almost anything unless you are like Google or Apple.
But if you don't work there, then you are basically safe to do almost anything with it.
Suggestion: if you feel the need to share some LLM output with HN, prefix it with a human-written sentence or two explaining why it's worth reading.
Assuming zero copy, it should be loadable on around 170GB.
Reference: https://www.hopsworks.ai/dictionary/llms-large-language-mode...
Perhaps it would have been more obvious that I actually ran the model and measured it if people hadn't flagged the self-reply with the model output.
If you want to try it with Ollama on macOS (keeping in mind you'll need the new Mac Studio with 192GB of memory) this command will work:
ollama run falcon:180b
Right now it gets about ~5 t/s.. there's quite a bit of work being done to improve inference speeds like speculative sampling which uses a smaller model for a subset of tokensQuantization like ExLlama's EX2 or even llama.cpp's k-quants is already far from brute force: https://github.com/turboderp/exllamav2#exl2-quantization
> The format allows for mixing quantization levels within a model to achieve any average bitrate between 2 and 8 bits per weight. Moreover, it's possible to apply multiple quantization levels to each linear layer, producing something akin to sparse quantization wherein more important weights (columns) are quantized with more bits. The same remapping trick that lets ExLlama work efficiently with act-order models allows this mixing of formats to happen with little to no impact on performance. Parameter selection is done automatically by quantizing each matrix multiple times, measuring the quantization error (with respect to the chosen calibration data) for each of a number of possible settings, per layer. Finally, a combination is chosen that minimizes the maximum quantization error over the entire model while meeting a target average bitrate.
Llama.cpp is also working on a feature that let's a small model "guess" the output of a big model which then "checks" it for correctness. This is more of a performance feature, but you could also arrange it to accelerate a big model on a small GPU.
It's not clear to me whether compilations of model weights are really copyrightable in most jurisdictions. Even if they were, enforceability is highly questionable. Anyway downloading is not the same as uploading; the recipient is not the one making the copy.
So long as you are not redistributing the model weights but just using those model weights to do something, I don't see how you're violating copyright. If you pirate a compiler and distribute your own software compiled with it--is that compiled software a derivative work of the compiler? Is my proprietary software a derivative work of g++ because I used it to compile my software?
Even providing a direct API to inference with or finetuning a model should be OK. I can't see how the output of inference could be a derivative work of the model, which is only part of the input along with the hyperparameters, seed value, potential fine-tuning data, and the prompt. If, alternatively, it was, then the output of training (i.e. the model itself) is surely a derivative work of the input--including my data they slurped up!
On the other hand, redistributing finetuned models might need to be done by weight diffs, as those probably are derived works.
[0]: e.g. b8287ebfa04f879b048d4d4404108cf3e8014352 and 14ec42c98d0b12f5b12f5b3d670cdf4915c1feaf
Anyway, running copyrighted software without a valid license is clear cut unlawful, at least in the EU. That's because running the application requires copying it (into RAM).
No, really, the law is that pedantic. "Loading" is a restricted act under Article 4 of the Computer Programs Directive.