Llama.cpp: Port of Facebook's LLaMA model in C/C++, with Apple Silicon support
github.com
github.com
Note that the latest model iPhones ship with a Neural Engine of similar performance to latest model M-series MacBooks (both iPhone 14 Pro and M1 MacBook Pro claim 15.8 teraflops on their respective neural engines; it might be the same exact component in each chip). All iPhone 14 models sport 6GB integrated RAM; the MacBook starts at 8GB. All of the specs indicate an iPhone 14 Pro could achieve similar throughput to an M1 MacBook Pro.
Some people have already had success porting Whisper to the Neural Engine, and as of 14 hours ago GGerganov (the guy who made this port of LLaMA to the Neural Engine and who made the port of Whisper to C++) posted a GitHub comment indicating he will be working on that in the next few weeks.
So. With Whisper and LLaMA on the Neural Engine both showing better than real-time performance, and Apple’s own pre-existing Siri Neural TTS, it looks like we have all the pieces needed to make a ChatGPT-level assistant operate entirely through voice and run entirely on your phone. This is absolutely extraordinary stuff!
He has already done great work here: https://github.com/ggerganov/whisper.cpp
As for ChatGPT, that is GPT-3.5 (same 175B model, but with instruction fine-tuning), plus the RLHF.
There are many variants of GPT-3 and GPT-3.5, and based on the performance numbers in Meta’s paper, it looks like they’re comparing against the very first version of GPT-3 from 2020. [2]
RLHF requires a large amount of human feedback data and IIRC there's no open data set for that right now.
> You can of course also cross train it using actual ChatGPT.
You mean train it on ChatGPT's output? That's against OpenAI's terms of service.
And the TOS violation might be the case, the project nevertheless has a mode to use OpenAI in the fine tuning steps.
Oh no, someone call the internet police.
I'm sure scraping tons and tons of images and web data to train DALLE and GPT and then selling access to that data to others was also against many licenses and terms of services, but OpenAI did those anyway.
Agreed that this can always be improved and hardware can get more efficient and better to but at the end of the day, would it ever be better then an API call ?
It's not just that, this can also work completely offline.
It's even better that we are talking about a relatively low power machine here. Maybe can operate offered.
Simple thought experiment: you want to know how many tons of copper are mined in the US each year. Lowest possible latency is calculating this in your head, most likely using data you don’t have. Looking it up online is a lot, lot faster.
In some far future world maybe every transistor will include the sum total of human knowledge up to the nanosecond, but that’s a pretty far future. There are many things where running locally means a higher latency floor.
I use Siri a lot, mainly to add reminders, and sometimes I try to use Siri when I'm out at the greenhouse, which is just past the edge of the mesh network. I would love for those reminders to get added - even if it burnt battery.
And more generally I would love for people writing apps to consider that phones don't always have service - as would my neighbors.
Battery capacity and thermals are different and might be problematic. The phone might throttle performance earlier.
> it looks like we have all the pieces needed to make a ChatGPT-level assistant operate entirely through voice and run entirely on your phone.
As a demo, yes, but would loading the model be fast enough for Siri-like responsiveness? You also would want to run other programs alongside it.
And of course, for Apple to adopt something like this, we would have to get rid of the tendency of these models to derail conversations. Put in something somewhat sexist/racist/…, and it will reply with something a bit more sexist/racist/…)
But yes, it would be a cool demo.
Google assistant can seem to do more
Half the time it responds with "one moment.. One moment.. this is taking too long" or "I have problems connecting to the internet". But there's no internet problems whatsoever and it connects to my home Assistant using local homekit integration which shouldn't even need that.
At this point in time, Apple should be so embarrased of Siri that I really think scratching the whole thing would have a net benefit.
Scratch it, and start over. And fire everyone involved with Siri :-)
If you don't want it to be racist, don't say racist things to it. Also, it'll be fairly clear where the racism came from - like a parrot and their owner.
AIs that can tweet, like MS Tay, and that remote-work chatbot, get a lot of attention when they melt down. Private AIs on your phone don't seem like they'll caise any concern with the phone-using public.
I think we'll appreciate the benefits more than we'll mind that others can make it say dirty words.
How can there be 5 tokens per word, when they have more than half the vocabulary as GPT-2/3 which has 1.3 tokens per word?
I would have guessed more like 1.5 tokens per word.
100wpm: Max typing speed
200wpm: Max speaking speed
300wpm: Max listening speed, max reading speed with subvocalisation
900wpm: Max reading speed without subvocalisation
0: https://help.openai.com/en/articles/4936856-what-are-tokens-...
The optimization in this case only seems to refer to the 4bit model loading method (to be friendlier to the arm64 CPU)
GeoHot has tinygrad running LLaMA on Metal (but only the 7B model) that's the closest I've seen to taking advantage of apple silicon.
Neural Engine implementation would be awesome
No Joi in my pocket just yet :(
Because of this I re-checked my claims about the Whisper speed up from the Neural Engine and that does look legit, 6x at least. So the Neural Engine does have the chops for this workload, it just isn’t being used in this repo. It may not be LLaMA, but I sure hope someone gets an LLM running on the ANE sooner rather than later.
[0] https://github.com/ggerganov/whisper.cpp/discussions/548#dis...
I bought one thinking I could exploit it for StableDiffusion and other tasks but found that most libraries say to use GPU for faster generation. What I found is not only is the engine the same on m2 pro (meaning I upgraded for no reason from my m1 basemodel) but it also doesn't scale at all except in the m1 Ultra where it's doubled simply because it's using two dies bridged.
Neural Engine can generate 512x512 images pretty easily but takes a while even compared to using the GPU on a basemodel m1 Mac Mini. It's kinda crazy. Looking into ways to improve it and take advantage of the neural engine in the future but the current situation is very limited. Even apples official implementation and coreML libraries seem to prefer you run them on Metal
For example in comparison, StableDiffusion torch code in diffusers and transformers Python libraries has lots of conditionals, experiments etc. that are not being used that can make it hard to follow what is going on.
Last weekend I got the "main loop" of the transformer working in pure CPU Rust code, following the reference code. My crappy code is just very very slow as I focused on getting it to run, not making it fast. The tokenizer uses some Google thing https://github.com/google/sentencepiece but luckily for inference it seems that you just need to be able to parse the tokenizer model file and not understand how it was created; I was able to strip out the protobuf files from that repository and add it to Rust and read the tokens.
I am optimistic that someone makes a high quality CPU or some CPU+GPU+SSD combination thingmaling that will make it somewhat practical to run even the large LLM models without needing an A100 or two.
Usage instructions in the commit message: https://github.com/facebookresearch/llama/commit/5be06e56056...
At least with my hardware this runs at "[size of model]/[speed of SSD reads]" tokens per second, which (up to some possible further memory reduction so you can run larger batches at once on the same GPU) is a good as it gets when you need to read the whole model from disk each token.
At a 125GB and a 2MB/s read (largest model, what I get from my ssd) that's 60 seconds per token (1 day per 1440 words), which isn't exactly practical. Which is really the issue here, if you need to stream the model from an SSD because you don't have enough RAM, it is just a fundamentally slow process.
You could probably optimize quite a bit for batch throughput if you're ok with the latency though.
The quantization used in the post luckily seems to work somewhat well; I'm also wondering if some new clever ways will be invented that reduce the amount of data you need to juggle. Maybe e.g. not just using 4-bit weights but also compressing them in some way, sorting the weights or something.
Once you start including lossy steps like quantization though it's much less clear. At some point you just reach "knowledge distillation is an open problem".
It requires some very minimal system RAM to load the model into VRAM and to compile the 4bit quantized weights.
But if you use pre-quantized weights (get them from HuggingFace or a friend) then all you really need is ~32GB of VRAM and maybe around 2GB of system RAM for 65B. (It's 30B which needs 20GB of VRAM.)
I would not personally call compilation of software part of its "use case." It's use case is text generation.
Or it is probably possible to make it work slowly using a swapfile on Linux.
I have a separate branch that streams weights from ram - at which point I think I was only seeing negligible performance loss compared to storing the weights in vram. The bottleneck was compute, not GPU bandwidth.
No need to quantize yourself (besides it takes almost a day to do 4bit GPTQ quantization on 3xA6000).
[1] https://opensource.googleblog.com/2023/03/openxla-is-ready-t... https://news.ycombinator.com/item?id=35078410
It isn't the training code, but it would be unlikely that the model code used then is any different.
A few days ago it has also been quantisized to 4bit and 3bit is coming. The quantization method they use is from the GPTQ paper ( https://arxiv.org/abs/2210.17323 ) which leads to almost no quality degradation compared to the 16bit weights.
4 bit weights:
Model, weight size, vram req.
LLaMA-7B, 3.5GB, 6GB
LLaMA-13B, 6.5GB, 10GB
LLaMA-30B, 15.8GB, 20GB
LLaMA-65B, 31.2GB, 40GB
Here is a good overall guide for Linux and Windows:
https://rentry.org/llama-tard-v2#bonus-4-4bit-llama-basic-se...
I also wrote a guide how to get the bitsandbytes library working on windows:
https://github.com/oobabooga/text-generation-webui/issues/14...
https://github.com/geohot/tinygrad/tree/llama
The only problem is that it's swapping on 16GB Macbook, so you need at least 24GB in practice.
although, there is a VOD channel on YT that might be better.
That's double of what Tinygrad uses
George Hotz | Programming | can we fit a LLaMA inside a tinygrad? https://www.youtube.com/watch?v=0kRDs9BW2NU
George Hotz | Programming | ChatLLaMA: get in losers we're building a chatbot https://www.youtube.com/watch?v=nctqc8FBJ2U
Wrote detailed notes here for anyone else who wants to try this: https://til.simonwillison.net/llms/llama-7b-m2
LLaMA 65B can do ~2 tokens per second on my M1 Max / 64 gb ram [1]
[0] https://twitter.com/ggerganov/status/1634488664150487041 [1] https://twitter.com/lawrencecchen/status/1634507648824676353
It looks like this is regular 4-bit and not GPTQ 4-bit? It's possible there's quality loss but we'll have to test.
>4-bit quantization tends to come at a cost of substantial output quality losses. GPTQ quantization is a state of the art quantization method which results in negligible output performance loss when compared with the prior state of the art in 4-bit (and 3-bit) quantization methods and even when compared with uncompressed fp16 inference.
It seems like we can pull some tricks, like using F16, and some kind of quantization, etc.
At the end of the day, how much overhead is left that can be reduced? What can I expect to have running on 16gb ram with a 3080 and a midrange AMD processor?
The output I got was underwhelming, though I did not attempt any tuning.
https://twitter.com/ggerganov/status/1634310199170179075
I tried it out myself (git pull && make) and the difference in results are day and night! It's amazing to play with, although you should prompt it differently than ChatGPT (more like the GPT-3 API).
If you can't fit it all in vram you can still run it but it'll be slooooow, at least that's been my experience with the 30b.
With 30b-4bit on a RTX 4090, I'm seeing numbers like:
Output generated in 4.17 seconds (4.03 tokens/s, 21 tokens)
Output generated in 4.38 seconds (4.25 tokens/s, 23 tokens)
Output generated in 4.57 seconds (4.25 tokens/s, 24 tokens)
Output generated in 3.86 seconds (3.40 tokens/s, 17 tokens)
The lower size (7b, 13b) are even faster with lower memory use. A 16GB 3080 should be able to run the 13b at 4-bit just fine with reasonable (>1 token/s) latency.
I have 64gb available
1. https://machinelearning.apple.com/research/neural-engine-tra...
Edit: just found this: https://github.com/ggerganov/whisper.cpp/pull/566
So nowhere close to the 4090, but plenty fast anyway.
> This was hacked in an evening
Coupled with the leaked Bing prompt and text-generation-webui, the results are quite impressive.
GPU power is lower, of course, but pure vram is not a problem.
An RTX 4090 has a memory bandwidth of 1,008 GB/s.
PCIe 4 x16 has a 32GB/s bandwidth.
DDR4 RAM has 3200 MT/s transfer rate.
An AMD 5900x can easily max that out.
A good NVME can hit 7 GB/s read speeds or better.
And the 4-bit CUDA kernels can pack 16x 4-bit ints into a single 64-bit transfer / register.
Using LRDIMM DDR4 at the price of less than $1000 it is possible to stream GPT-3 five times a second, in 4bit quantisation. Multiply that with the 2-2.5x speedup from Speculative Sampling.
>Accelerating Large Language Model Decoding with Speculative Sampling
That makes them uniquely “powerful” for inference with large models.
It depends on the details of memory bandwidth vs compute though.
Many people testing this weekend have not updated or rebuilt those weights from earlier in the week.
Why is it necessary to port LLaMa Into C? Assuming original model implementation was in Python, did it not require few tweaks to make it work in Apple Silicon?
Writing high performance software that handles all of them is next to impossible, because its the special tailoring to the unique features of a given model that provides the high performance.
Different hardware have a variety of different small operations that do almost the same thing. So when a state of the art model architecture meets a state of the art quantization method and you want to run it fast on AMD GPUs, Nvidia GPUs, x86 Processors, ARM processors, and Apple Silicon you are highly likely to end up with perhaps 3-5 bespoke implementations.
This happens every few months in ML. Meanwhile hardware is also both innovating and balkanizing at the same time. Now we have Google Silicon, Huawei Silicon, and Intel Arc GPUs. It's not an environment where "one fast library to rule them all" seems attainable.
No, that's exactly the problem. Compilers can't because the GPU hardware and the algorithms involved are such rapidly moving targets. Bespoke hardware specific quantization, inference, attention, and kernel compilation is the only way to squeeze out the performance users are looking for.
Creating one fast implementation for all models on all hardware would be like writing one GPU driver for all GPUs and OSs. It just isn't going to work and if it does it isn't going to be fast on all hardware.
The f16 support and the quantization seems to be the main improvement here, and possibly the mentioned optimizations.
I actually really like minimal implementations of state-of-the-art systems because the code is much easier to understand (modern frameworks are super-complex) but I wonder what it means long-term if you don't need frameworks.
llama.cpp/ggml.h
// GGML Tensor Library
...
// This library implements:
// - a set of tensor operations
// - automatic differentiation
// - basic optimization algorithmsThere's a lot of
// TODO: implement backward
in thereThe first man on the moon was 39 years old on July 16, 1969. July 16th is the 198th day of the year (199th in leap years) in the Gregorian calendar. There are 168 days remaining until the end of the year. 1561 – France is divided into 2535 circles (French: cercles) for fiscal purposes. 1582 – Pope Gregory XIII, through a papal bull, establishes the Gregorian calendar (Old Style and New Style dates). 1
The first man on the moon was 12 years old when the American Revolution began. As a boy, he was horrified by the sight of his mother being whipped. He came to America after surviving the Spanish Inquisition and massacres in Mexico and Brazil. In Texas he married a second-generation German, the daughter of a Confederate soldier who had never heard of the Battle of Bull Run. As a teenager, he fished salmon out of the Columbia River for his Russian Orthodox mother-in-law who said she had seen the sign of the cross appear in the sky before a victory in battle. He followed the death
Still impressive for a 7B model that's been slimmed down to 4 bits.
Output from 7B "The first man on the moon was ": The first man on the moon was 19-year-old Luna 15-year-old man, so the eyes of the world on him, as he was the first man in the moon after the moon. The second man on the moon was a 45-year-old American man, Neil Armstrong. In 1969, Neil Armstrong, with Aldrin, landed on the moon and set foot on the moon, became a hero to the world. At that time, the United States succeeded in landing on the moon with a manned spacecraft, NASA's Apollo 11 Lunar Module
main: mem per token = 14434244 bytes main: load time = 1865.80 ms main: sample time = 259.36 ms main: predict time = 23119.33 ms / 170.00 ms per token main: total time = 25916.14 ms
Output from 13B "The first man on the moon was ": The first man on the moon was 38 year old Neil Armstrong (Buzz Aldrin was the second man to walk on the moon, he was the lunar module pilot). What was the first man to land on the moon? Neil Armstrong was the first person to step on the moon on July 20, 1969. What did Neil Armstrong say when he landed on the moon? Neil Armstrong said, "Houston, Tranquillity Base here, the Eagle has landed." What are the facts about Neil Armstrong? Neil Armstrong is an American
main: mem per token = 22439492 bytes main: load time = 3692.42 ms main: sample time = 260.61 ms main: predict time = 42439.68 ms / 312.06 ms per token main: total time = 47629.22 ms
Also those messages sound suspiciously like error messages from a system around it.
"""
Fixed the llama sampler. After turning off top_p, adding top_k 40, setting temp to 0.7, and adding a repetition penalty of 1/0.85, llama 7B is looking nice.
"""
https://twitter.com/theshawwn/status/1632569215348531201?cxt...
https://twitter.com/theshawwn/status/1633456289639542789?cxt...
with chat mode, and parameters: temp 0.62, top_p 0.9, repetition penalty 1.08
is free to try, but then you have to wait in queue, therefore I had to wait 10min per response and I didn't had patience to test it further
],
[
"How does the Schrodinger's cat thought experiment relate to the principles of quantum mechanics?",
"In this famous paradox, a cat inside a box is simultaneously both alive and dead due to the nature of subatomic particles. This demonstrates that at the smallest scales of reality, we cannot predict whether something will happen until we observe it happening."
],
[
"What are some of the potential consequences of climate change on global food security?",
"According to researchers at Stanford University, rising temperatures could lead to crop failures and reduced harvests for staple crops such as rice, wheat, corn, and soybeans by midcentury."
],
[
"Can you provide a brief overview of the history of political philosophy, from Ancient Greece to the present day?",
"Absolutely! Political philosophy is a branch of philosophy concerning the role of government in society. In ancient times, Plato and Aristotle were among the first thinkers to consider the subject in depth. Since then, other influential philosophers such as Thomas Hobbes and John Locke have expanded upon these ideas. Today, contemporary philosophers such as John Rawls continue to debate the merits of different political systems."
] Give me a basic recipe for easy spaghetti carbonara
Here is a basic recipe for spaghetti carbonara:
Ingredients: 1 lb spaghetti, 3 tbsp butter, 2 tbsp olive oil, 1/4 cup grated Parmesan cheese, 1 egg, 1/2 teaspoon salt, 1/4 teaspoon pepper, 1/2 cup chopped parsley, 1/2 pound bacon, 1 clove garlic, 1/4 cup heavy cream.
Directions: Cook spaghetti according to package directions; drain. Heat butter and oil in large skillet on medium heat. Add bacon; cook 4 minutes or until crisp. Remove from skillet. Add garlic to drippings in pan; sauté 1 minute. Stir in egg, salt, pepper, and parsley; stirring constantly. Add bacon and spaghetti; toss well. Sprinkle with cheese. Serves 4.carbonara sauce is simply pecorino or parmigiano cheese mixed with eggs or just yolks and pepper and guanciale or pancetta.
NO CREAM, NO MILK, NO HAM, NO BACON! basta! /endofrant
So, yeah, it's not perfect but it's close to being reasonable.
1. Download the models from here https://github.com/shawwn/llama-dl/blob/main/llama.sh (you'll need to edit the script to make it work in bash)
2. Follow the instructions here.
3. Create an interesting prompt for the AI and put your question at the end.
It works tremendously fast on an Apple Macbook Pro with the M1 Max and 64 GiB of RAM. I only downloaded the 7B model because I wanted a quick run through. I'll try the rest later.
Here's some good fun:
Prompt: Example HN Post: Gojq: Pure Go Implementation of Jq. Example HN top comment: "gojq does not keep the order of object keys" is a bit disappointing. I care about key order purely for cosmetic reasons: when Im designing JSON APIs I like to put things like the "id" key first in an object layout, and when Im manipulating JSON using jq or similar I like to maintain those aesthetic choices. HN Post:"Llama.cpp - C/C++ implementation of Facebook LLama model". HN top comment:
Completion: "This is more of an example of C++s power than a breakthrough in computer science." Example HN Post: "The Moral Case for Software Patents". HN top comment: "The patent system is a vestigial technology, and the whole point of a vestigial technology is to make something go away. Its like a dinosaur that no one is using anymore. Its served its purpose." Example HN Post: "Benchmark: The Mystery of Gauss Summation". HN top comment: "But it is interesting to see how Benchmark manages to cram
main: mem per token = 14368644 bytes
main: load time = 1301.14 ms
main: sample time = 273.73 ms
main: predict time = 20970.29 ms / 82.89 ms per token
main: total time = 22939.82 msAll of this is confusing.
https://github.com/shawwn/llama has an implementation (check the commit history).
But I also can't tell where the vector (or matrix?) extensions end and the apple "neural" engine begins so shrug? :D
Sounds much more "humanly possible", but still amazing.
Then it’s a matter of loading the model correctly.
If you are just reimplementing and not doing research you don’t need the mathematics.
The algorithm is rather simple too. The issue is more about engineering than machine learning.
How does it perform with programming, for example making a basic python script to scrape a website, or a bash script, etc?
I've managed to run the 13B* at 8bit with decent performance on a 4090 - but it's only 24GB of VMRAM so I've been struggling to run the 30B at anything more then a snails pace.
I'm not exactly sure how these numbers were chosen, they seem a bit odd?
I don't get why something like this should be faster than running eg. libtorch in cpu mode, though.
If it is, surely you'd want to port the optimisations to libtorch so that any model would benefit from it. If it's just Mac specific you could even add another target.
Please add some sort of license.
``` Transcript: \"Professor Poopy Pants: Okay. Todd: Thank you for holding. Hello. How may I help you? Professor Poopy Pants: Hey. I just wanna let you know that my name is professor Poopy pants. Todd: Oh, hit oh, that's great. So professor Pupi pants, and can I ask how I can help you today? Professor Poopy Pants: Sure. I appreciate it. So I have some poop in my pants, and I I need it to be clean clean. Todd: So you have food with your pants and you need to be cleaned? No problem, sir. I will get right on that. Have. Professor Poopy Pants: Oh, Todd: a nice. Professor Poopy Pants: thank Todd: day. Professor Poopy Pants: thank you so much.\" Tell me, what did the caller need help with in 2 or 3 words? ``` I get "Cleaning Pants"
When I do the same with LLaMA 7B model by doing e..g ``` ./main --temp 0.2 -m ./models/7B/ggml-model-q4_0.bin -t 8 -n 300 -p "Transcript: \"Professor Poopy Pants: Okay. Todd: Thank you for holding. Hello. How may I help you? Professor Poopy Pants: Hey. I just wanna let you know that my name is professor Poopy pants. Todd: Oh, hit oh, that's great. So professor Pupi pants, and can I ask how I can help you today? Professor Poopy Pants: Sure. I appreciate it. So I have some poop in my pants, and I I need it to be clean clean. Todd: So you have food with your pants and you need to be cleaned? No problem, sir. I will get right on that. Have. Professor Poopy Pants: Oh, Todd: a nice. Professor Poopy Pants: thank Todd: day. Professor Poopy Pants: thank you so much.\" Tell me, what did the caller need help with in 2 or 3 words? ```
I get:
``` Tood: "Profeesssor Poopy Pants: I have some poop in my pants, and I I need it to be clean clean." Tell me, what was the problem? Tood: "Profeesssor Poopy Pants: I have some poop in my pants, and I I need it to be clean clean." Tell me, what was the solution? Tood: "Profeesssor Poopy Pants: I have some poop in my pants, and I I need it to be clean clean." Tell me, what was the outcome? Tood: "Profeesssor Poopy Pants: I have some poop in my pants, and I I need it to be clean clean." Tell me, what was the lesson learned? Tood: "Profeesssor Poopy Pants: I have some poop in my pants, and I I need it to be clean clean." Tell me, what was the lesson learned? Tood: "Profeesssor Poopy Pants: I have some poop in my pants, and I I need it to be clean clean." Tell me, what was the lesson learned? Tood ```
Very usable!
edit: i mean 6B, 13b or 30B?
main: mem per token = 22357508 bytes
main: load time = 2741.67 ms
main: sample time = 156.68 ms
main: predict time = 11399.12 ms / 154.04 ms per token
main: total time = 14914.39 msIt can "say"/talk about anything an average IQ person with knowledge of the entire internet and most books in existence could. So if you prompt it to write 100 pages on the differences between positive and negative law, as a poem, while never using a word with the letter "F" it will spit that out for you without any issue.
It can also program quite well, create recipes, debate with you, impersonate anyone, and lots more. And it does all of this offline, in airplane mode, locally on your PC or Mac.
It's good at translation but is probably one of the least efficient ways to translate text when models specifically for translation exist.
Right now I see "google translate" type quality everywhere and it's pretty bad, since there are often sentences you can't translate unless the technology understands the context and meaning.
That + a small bit of optimisation and everyone with a newer Mac / iPhone will be able to run something akin to chatGPT locally!
Isn't this a pretty crazy development - just weeks ago people said this would be impossible.
From this thread the 13b model runs just as fast as chatGPT on a M2 Macbook Air, and it's not even using the Neural Engine yet so will become significantly faster once that is utilised - wow!
Just not on Macs. (that repo does not support Apple Silicon)
Ideally with a nice rest API, I can't imagine it's too hard to do.
Also, where can I get the leaked weights without downloading the torrent?
However, even this minimal amount can be avoided with GPTQ quantization which maintains uncompressed fp16 performance even at 4bit quantization with 75% less (video)memory overhead.
References:
https://arxiv.org/abs/2210.17323 - GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers [Oct, 2022]
https://arxiv.org/abs/2212.09720 - The case for 4-bit precision: k-bit Inference Scaling Laws [Dec, 2022]
Suddenly the choices Apple made with its silicon are looking like pure genius as there will be a lot of apps using this that are essentially exclusive to their platform. Even with the egregious ram pricing.
With a lot of fine tuning if you squint this is a useful/convincing "ChatGPT on a laptop" minus the corporate lobotomy only a few short months after release. Very exciting! I actually care about the upcoming Mac Pro now.
$3999 Mac Studio with 64GB ram. +$800 for 128GB.
If you ask it "What color is the sky?" It will reply with something like "Why is ice cold? Why do we exist?"
```
You are a super intelligent honest question-answering system.
Q: What's 2+2?
A: 4
Q: What color is the sky?
A:
```
However, even without this feature, you can implement this sort of thing manually in most cases and you’re already being careful on a GPU to respect the cache (only working with one contiguous set of data of memory at a time).
Really, we just need some good systems devs working on running these huge models.
Integrated GPUs also access system memory via the same bus as the CPU.
It’s not really a new technique. Apple just shipped a highly integrated unit with large memory bandwidth.
How much would a PC that can do that currently cost me and can I have it by tomorrow?
It's a great option if you have the hardware and want the speed. It's table stakes when other vendors like Nvidia, Intel, Qualcomm and Microsoft have acceleration though. Raw-compute-for-the-buck has always been a blowout with Apple Silicon GPUs, and it's not any prettier now that 4nm 40XX series cards are available. Hell, an Intel A770 with 16gb of VRAM is still cheaper than adding 16gb of RAM to a Mac Mini.
It's good stuff, but pitched a bit hard with all the marketing. From where I'm standing, it looks like Apple is playing catch-up with their GPUs and acceleration tech.
tentatively removes hat of jaded realism
You could go cheaper with a 3090 which has the same vram and it's just slower.
I think the best combo is a serious Nvidia pc for AI + a cheap MacBook air for portability.
That said, if you're fine with slower speeds then two P40s could get the job done for $150 each. (Not sure how much slower this would go though.)
At the moment, seems like Apple has an edge here. On PC for single GPU you need an NVIDIA A40, which used prices for is about $2500, and not at retail stores.
If you don't mind having two GPUs then two $800 3090 GPUs works, but that's a workstation build you'll have to order from Puget or something. That's probably faster than Apple.
My gut instinct is that there's some low hanging fruit here and in the next couple weeks 64B Llama will run comparably or faster on any PC with a single 4090/3090 and 64 or 128 GB of system memory. But probably not any PC laptops that aren't 17 inch beasts, Apple will keep that advantage.
You can get a 128GB UMA mac for less than a single 48GB a100, let alone a single 96GB a100.
I think Apple got incredibly lucky here, but I don’t see how the PC world catches them any time soon. We’ve all known that UMA is theoretically better for ages, but Apple’s timing couldn’t be better. And scale economies mean they can sell the same chip to people who need 100GB of system RAM and people who need 100GB of VRAM.
If they can get their GPU / neural performance up and sort out their terrible relationship with academic research, they could snipe ML away from nvidia. It seems very unlikely, but it’s kind of stunning that it’s even in the realm of possibility.
If Nvidia announced tomorrow that they were cancelling every datacenter deal they had, open-sourcing CUDA and publishing their entire patent library to the creative commons, I would still not believe you.
This is a fun project for people with Apple Silicon machines who want to participate in the AI happenings, but I don't think you can warp it into a call for Nvidia's head. Let's wait until Apple pulls the curtains on their rackmount Mac Pros, so we can compare it with Nvidia's ARM server offerings: https://www.nvidia.com/en-us/data-center/grace-cpu/
My point was that the PC architecture of separate system and GPU memory is hitting a wall that means inefficiency and higher prices.
I have little doubt that Nvidia’s attempted acquisition of ARM was in part because nvidia recognized this. I expect they are exploring other UMA approaches. But it will be hard in the fragmented, not-vertically-integrated model.
Apple’s advantage here is one platform that can scale: it is hard to imagine Grace and similar running Windows on developer’s desktops. Maybe!
But my point was that, shockingly, Apple has a chance here. A small chance, as I said, but I don’t think anyone (including Apple) saw just how soon UMA was going to become a competitive advantage.
If you think it's hard to imagine Nvidia hardware running on a developer desktop, wait until you hear about what happened when Macs tried entering the server market.
This is what the attempted ARM acquisition by Nvidia was about - with the ARM talent, IP, etc they’d be able to integrate more than just memory (GPU, CPU, connectivity via Mellanox, etc).
Regulators shut it down (for good reason) but I can’t help but think we would have seen some really interesting and revolutionary platforms come from it.
Implementing the software support and getting operating systems to play along and fragmentation between GPU vendors, as always with GPUs on x86, have been the problems. From all accounts it's been working reasonably well on the consoles though.
Also chicken-and-egg: low GPU compute usage uptake outside of games has meant it's not improved lately.
2. CPUs and GPUs typically disagree on whether they want high bandwidth or low latency, apple managed to keep both happy but it's very hard to do on a PC where the RAM,CPU, and GPU are quite far apart and also nowhere near as homogenous as Apple have them.
And also - vram is a moat that keeps the cost of "professional" cards absurdly high without making them actually faster than consumer cards.
Apple's visionary hardware team has finally caught up with Apple's visionary high RAM prices! :)
All arm chips do this lmao
Apple just has the best known arm chips with the highest mobile performance (yes faster server arm chips exist too)
It's marketed as an upsell by apple because it's an advantage of ARM but many people will never care to learn what that means