Using LLaMA with M1 Mac and Python 3.11
dev.l1x.be
dev.l1x.be
Additionally, llama.cpp works fine with 10 y.o hardware that supports AVX2.
I'm running llama.cpp right now on an ancient Intel i5 2013 MacBook with only 2 cores and 8 GB RAM - 7B 4bit model loads in 8 seconds to 4.2 GB RAM and gives 600 ms per token.
btw: anyone knows how to disable swap per process in macOS ? even though there is enough free RAM, sometimes macOS decides to use swap instead.
Edit: Thank you, @diimdeep!
To get running, just follow these steps https://github.com/ggerganov/llama.cpp/#usage
As @rnosov notes elsewhere in the thread, this post has a workaround for the PyTorch issue with Python 3.11, which is why the "and Python 3.11" qualification is there.
That takes at most a minute to run, but once converted you'll never need to run it again. Actual llama.cpp model inference uses compiled C++ code with no Python involved at all.
Well, on my macOS Ventura 13.2.1 install, /usr/bin/python3 is Python 3.9.6, which may be too old?
But also, my custom installed python3 via homebrew is not 3.11 either. My /opt/homebrew/bin/python3 is Python 3.10.9
MacBook Pro M1
I’m running 13B on MacBook Air M2 quite easily. Will try 30B but probably won’t be able to keep my browser open :/
shameless plug: https://mobile.twitter.com/tomprimozic/status/16348774773100...
It definitely runs. It uses almost 20GB of RAM so I had to exit my browser and VS Code to keep the memory usage down.
But it produces completely garbled output. Either there's a bug in the program, or the tokens are different to 13B model, or I performed the conversion wrong, or the 4bit quantization breaks it.
30B just says “dotnetdotnetdotnet…”
(venv) bherman@Rattata ~/llama.cpp$ ./main -m ./models/30B/ggml-model-q4_0.bin -
t 8 -n 128
main: seed = 1678666507
llama_model_load: loading model from './models/30B/ggml-model-q4_0.bin' - please wait ...
llama_model_load: n_vocab = 32000
llama_model_load: n_ctx = 512
llama_model_load: n_embd = 6656
llama_model_load: n_mult = 256
llama_model_load: n_head = 52
llama_model_load: n_layer = 60
llama_model_load: n_rot = 128
llama_model_load: f16 = 2
llama_model_load: n_ff = 17920
llama_model_load: n_parts = 4
llama_model_load: ggml ctx size = 20951.50 MB
llama_model_load: memory_size = 1560.00 MB, n_mem = 30720
llama_model_load: loading model part 1/4 from './models/30B/ggml-model-q4_0.bin'
llama_model_load: ................................................................... done
llama_model_load: model size = 4850.14 MB / num tensors = 543
llama_model_load: loading model part 2/4 from './models/30B/ggml-model-q4_0.bin.1'
llama_model_load: ................................................................... done
llama_model_load: model size = 4850.14 MB / num tensors = 543
llama_model_load: loading model part 3/4 from './models/30B/ggml-model-q4_0.bin.2'
llama_model_load: ................................................................... done
llama_model_load: model size = 4850.14 MB / num tensors = 543
llama_model_load: loading model part 4/4 from './models/30B/ggml-model-q4_0.bin.3'
llama_model_load: ................................................................... done
llama_model_load: model size = 4850.14 MB / num tensors = 543
main: prompt: 'When'
main: number of tokens in prompt = 2
1 -> ''
10401 -> 'When'
sampling parameters: temp = 0.800000, top_k = 40, top_p = 0.950000
When you need the help of an Auto Locksmith Kirtlington look no further than our team of experts who are always on call 24 hours a day 365 days a year.
We have a team of auto locksmiths on call in Kirtlington 24 hours a day 365 days a year to help with any auto locksmith emergency you may find yourself in, whether it be repairing an broken omega lock, reprogramming your car transponder keys, replacing a lacking vehicle key or limiting chipped car fobs, our team of auto lock
main: mem per token = 43387780 bytes
main: load time = 35493.44 ms
main: sample time = 281.98 ms
main: predict time = 34094.89 ms / 264.30 ms per token
main: total time = 74651.21 ms pip3 install --pre torch torchvision --extra-index-url https://download.pytorch.org/whl/nightly/cpu
That's the reason I stuck with Python 3.10 in my write-up for doing this: https://til.simonwillison.net/llms/llama-7b-m2 mamba create -n llama python==3.10 pytorch sentencepiece> Don’t get distracted by guys who are already out of your league; focus on the ones that have some hope for getting into it with them...even though they might not be there yet! Dont Forget To Sign Up and Watch our
(No, I didn't cut off the end. That's just how it stopped.) Anyway, makes it seem like, whatever their training corpus was, it deffo included scraping a bunch of social media influencers.
I'm currently running the 65B model just fine. It is a rather surreal experience, a ghost in my shell indeed.
As an aside, I'm seeing an interesting behaviour on the `-t` threads flag. I originally expected that this was similar to `make -j` flag where it controls the number of parallel threads but the total computation done would be the same. What I'm seeing is that this seems to change the fidelity of the output. At `-t 8` it has the fastest output presumably since that is the number of performance cores my M2 Max has. But up to `-t 12` the output fidelity increases, even though the output drastically slows down. I have 8 perf and 4 efficiency cores, so that makes superficial sense. At `-t 13` onwards, the performance exponentially decreases to the point that I effectively no longer have output.
I'm sure there are potential uses but training your own LLM would probably be more meaningfully useful versus running someone else's trained model, which is what this is.
> enough RAM
Because boring consumer laptops are of course known for their copious amounts of expandable RAM and not for having one socket fitted with the minimum amount possible.
> /main -m ~/Downloads/llama/7B/ggml-model-q4_0.bin -t 6 -n 256 -p 'The first man on the moon was '
The first man on the moon was 38 years old.
And that's when we were ready to land a ship of our own crew in outer space again, as opposed to just sending out probes or things like Skylab which is only designed for one trip and then they have to be de-orbited into some random spot on earth somewhere (not even hitting the water)
Warren Buffet has donated over $20 billion since 1978. His net worth today stands at more than a half trillion dollars ($53 Billiard). He's currently living in Omaha, NE as opposed to his earlier home of New York City/Berkshire Mountains area and he still lives like nothing changed except for being able to spend $20 billion.
Social Security is now paying out more than it collects because people are dying... That means that we're living longer past when Social security was supposed to run dry (65) [end of text]The main change needed seems to be InstructGPT style tuning (https://openai.com/research/instruction-following)
You have to lean on much older prompt engineering tricks - there are a few initial tips in the LLaMA FAQ here: https://github.com/facebookresearch/llama/blob/main/FAQ.md#2...
./main -m ./models/7B/ggml-model-q4_0.bin \
--top_p 2 --top_k 40 \
--repeat_penalty 1.176 \
--temp 0.7
-p 'async fn download_url(url: &str)'
async fn download_url(url: &str) -> io::Result<String> {
let url = URL(string_value=url);
if let Some(err) = url.verify() {} // nope, just skip the downloading part
else match err == None { // works now
true => Ok(String::from(match url.open("get")?{
|res| res.ok().expect_str(&url)?,
|err: io::Error| Err(io::ErrorKind(uint16_t::MAX as u8))),
false => Err(io::Errorrepeat_penalty is not an option.
./main -m ./models/7B/ggml-model-q4_0.bin \
--top_p 2 --top_k 40 \
--repeat_penalty 1.176 \
--temp 0.7
-p 'To seduce a woman, you first have to'
output: import numpy as np
from scipy.linalg import norm, LinAlgError
np.random.seed(10)
x = -2\*norm(LinAlgError())[0] # error message is too long for command line use
print x [end of text]Meta’s LLaMA v. GPT-3 comparisons are to OpenAI’s 2020 release of GPT-3, but there’s been significant progress since then.
Also, a reminder to folks that this model is not conversationally trained and won't behave like ChatGPT; it cannot take directions.
ChatGPT is GPT-3.5 Turbo.
3.5 Turbo is the ChatGPT model: it's cheaper (1/10th the price), faster and has a bunch of extra RLHF training to make it work well as a safe and usable chatbot.
https://openai.com/blog/introducing-chatgpt-and-whisper-apis introduced the turbo model.
MY MISTAKE: 002 and 003 are 3.5, but 001 looks to have pre-dated the InstructGPT work.
Might try other methods to do 30B, or switch to my M1 Macbook if that's useful (as it said here). Don't have an immediate need for it, just futzing with it currently.
I should note that web link is to software for a gradio text generation web UI, reminiscent of Automatic1111.
Maybe it could even run the 30b model?
1. There is no thoughput benefit to running on GPU unless you can fit all the weights in VRAM. Otherwise the moving the weights eats up any benefit you can get from the faster compute.
2. The quantized models do worse than non-quantized smaller models, so currently they aren't worth using for much use cases. My hope is that more sophisticated quantization methods (like GPTQ) will resolve this.
3. Much like using raw GPT-3, you need to put a lot of thought into your prompts. You can really tell it hasn't been 'aligned' or whatever the kids are calling it these days.
Assuming a sensible, somewhat linear layout using mmap to map the weights would give you the ability to load a lot in memory, with potentially a fairly minimal page-in overhead
https://twitter.com/lawrencecchen/status/1634507648824676353
I have had good luck in the past with Apple's TensorFlow tools for M1 for building my own models.
(On Linux with a 5800X and 64GB of RAM)
"The search for extraterrestrial life will most likely conclude adv provinß wojewłożGener Wikipédia Świirc Patrickvidcido protectsobDra"
How can it have all of this information packed into 4GB? I can't even imagine it being only 240GB.
These models have an unthinkable amount of information living in them.
For example, I assume LLaMa was trained on English Wikipedia data (it just makes sense). So let me try to prompt the 13B parameter model (which is 25 GiB to download, and 3.8 GiB after quantization) with "According to Wikipedia, the main causes of the French revolution are". It will give me the following continuations:
Attempt 1:
1) Social injustice and inequality;2)...
i need to do a report about the french revolution for my history class. so far i have these three questions:... [end of text]
Attempt 2: 1. The Enlightenment and its new ideas in philosophy had a great impact on France especially with regards their attitudes towards religion. There was also some political unrest due mainly from disagreements over how taxes were spent by different governments at that time....
What Were Three Causes of the French Revolution? The three causes for revolution are economic, social and governmental in nature... [end of text]
Attempt 3: 1) The Estates General was called by Louis XVI in order solve financial problems. However when they met there were three groups: Nobility who wanted a lot more power and money for themselves; Clergy (Bishops etc.) who also wanted more wealth and power but on top they didn’t want to pay taxes or be governed any other way then the old fashion “divine right” 3) The third group was called the Third Estate. They were made up of middle class citizens, peasants, artisans etc… These people had no political power and wanted more equality in society
While the Wikipedia article has a section on causes that starts with: The underlying causes of the French Revolution are usually attributed to the Ancien Régime's failure to manage social and economic inequality. Rapid population growth and the inability to adequately finance government debt resulted in economic depression, unemployment and high food prices. Combined with a regressive tax system and resistance to reform by the ruling elite, it resulted in a crisis Louis XVI proved unable to manage.
So the model is completely unable to reconstruct the data on which it was trained. It does have some vague association between the words of "French revolution", "causes", "inequality", "Louis XVI", "religion", "wealth", "power", and so on, so it can provide a vaguely-plausible continuation at least some of the time. But it's clear that a lot of information has been erased. CCNet [67%], C4 [15%], GitHub [4.5%], Wikipedia [4.5%], Books [4.5%], ArXiv [2.5%], Stack Exchange[2%]. The Wikipedia and Books domains include data in the following languages: bg, ca, cs, da, de, en, es, fr, hr, hu, it, nl, pl, pt, ro, ru, sl, sr, sv, ukEven a 128GB RAM M1 Ultra can’t run 65B unquantized.
Or they formatted the date yyyy/dd/mm but mistakenly wrote 08 instead of 03 for the month?
I just ran it on a plain old x86 servers with 64 cores and loads of RAM.
Works just fine. Apple H/W and Python version are completely irrelevant.
The first president of the USA was 57 years old when he assumed office (George Washington). Nowadays, the US electorate expects the new president to be more young at heart. President Donald Trump was 70 years old when he was inaugurated. In contrast to his predecessors, he is physically fit, healthy and active. And his fitness has been a prominent theme of his presidency. During the presidential campaign, he famously said he would be the “most active president ever” — a statement Trump has not yet achieved, but one that fits his approach to the office. His tweets demonstrate his physical activity.
Eh? Which bit is politically incorrect?I gave it a prompt containing the word of n and it actually ignored it but started talking about the Jews in terms that would make 4channers blush.
Not being reliant on a single entity is nice. I will accept not being on the bleeding edge of proprietary models and slower runs for the privacy and reliability of local execution.
I just wish these weren't all articles about how to run it on proprietary mac setups. I'm still waiting for the guides on how to run it on a real PC.
Mac __is__ a real PC.
The benefit is have a super-genius oracle in your pocket on-demand, without Microsoft or Amazon or anyone else eavesdropping on your use. Who wouldn't see the value in that?
In the coming age, this will be one of the few things that could possibly keep the nightmare dystopia at bay in my opinion.
But more importantly i want it uncensored. These tools are useful for deep conversation, which no longer exists online since many years ago