Llama2.c: Inference llama 2 in one file of pure C
github.com
github.com
- I wouldn't use anything higher than a 7B model if you want decent speed.
- Quantize to 4-bit to save RAM and run inference faster.
Speed will be around 15 tokens per second on CPU (tolerable), and 5-10x faster with a GPU.It's a shame the current Llama 2 jumps from 13B to 70B. In the past I tried running larger stuff by making a 32GB swap volume, but it's just impractically slow.
Also its really tricky to even build llama.cpp with a BLAS library, to make prompt ingestion less slow. The Oracle Linux OpenBLAS build isnt detected ootb, and it doesn't perform well compared to x86 for some reason.
LLVM/GCC have some kind of issue identifying the Ampere ARM architecture (march=native doesn't really work), so maybe this could be improved with the right compiler flags?
The OpenBLAS package was missing on ARM, along with some other dependencies I needed for compilation.
At the end of the day, even with many tweaks and custom compilation flags, the instance was averaging below 1 token/sec as a Kobold Horde host, which is below the threshold to even be allowed as a llm host.
I can see the ARM64 versions on the Ubuntu web package list, so... IDK what was going on?
On Oracle Linux, until I changed some env variables and lines in the makefile, the openblas build would "work," but it was actually silently failing and not using OpenBLAS.
That is going to be hard as the 7B model was trained on 2T tokens. Maybe if you heavily restrict the range in which the model should operate.
2. Better than tokens is to train on probability distributions (distillation) and trees of probability distributions
I did try a quick search for it. Found some interesting papers. The links to them are below in case anyone finds them interesting.
https://arxiv.org/abs/2212.11481
https://towardsdatascience.com/a-new-way-to-predict-probabil...
https://arxiv.org/pdf/1912.07913.pdf
https://dukespace.lib.duke.edu/dspace/bitstream/handle/10161...
ML algorithms are, at their core, not particularly complicated code. But they are still tricky code, because if you get them wrong you will find that you spent 500 GPU-years turning random numbers that cause the model to output gibberish into other random numbers that cause the model to output different yet semantically identical gibberish.
Writing them in a more abstract languages has advantages - like automatic differentiation. You could explicitly tell the computer how to compute the output and its derivative, or you could tell the computer how to compute the output, and let it also compute the derivative by itself.
Having all your weights in one object is also awfully convenient; you can write something like `weights -= error * deriv * learning_rate` instead of iterating over each individual weight (and a large model contains many different sets of weights, not just a single NxMxPxQ matrix)
This is good for the rapid iteration that ML research demands. However, once you have selected a model, I'm sure you can get performance advantages by coding it in a low level and eliminating inefficiencies. For example, you should be able to use the weight update equation from above by using fused multiply-accumulate, and the Python framework might not realize that.
https://pytorch.org/tutorials/advanced/cpp_frontend.html
In other words, you can absolutely use PyTorch without Python.
https://ggerganov.com/llama2.c/
Via his Twitter with ongoing thread: https://twitter.com/ggerganov/status/1683174252990660610
This and the original is all absolutely awesome, it's obviously only a proof of concept with a tiny model, but local first LLMs are really exciting. I particularly love the idea of being able to build webapps with local inference.
With optimisation, research into ways to make smaller models, partial downloads, and then the opportunity to use WebGPU we potentially have the start of an exciting new way to build privet local LLM based apps.
It's never going to be up to the same capabilities of hosted LLMs on massive clusters of top end GPUs, but there are so many use cases that this sort of thing will enable.
Once upon a time, there was a little girl named Lily. She loved to play outside in the park. One day, while she was playing, she saw a black bird flying in the sky. It was a beautiful bird with yellow wings.Lily ran to her friend, Timmy, and said, "Look, Timmy! A pretty bird!" Timmy smiled and said, "I see it! It's black and black."Suddenly, the sky turned dark and it started to rain. Lily and Timmy ran to a shelter and waited for the rain to stop. When it finally stopped, they ran back to Lily's house. They were happy to be safe and dry. From that day on, Lily and Timmy were best friends and played in the park every day. Once upon a time, in a small town, there was a big temple. Many people went to the temple to talk to each other. One day, a little boy named Tim went to the temple with his mom.Tim saw a pretty red ball at the temple. He asked his mom, "Can I have the ball, please?" His mom said, "Yes, you can, but we have to be polite his mommy washterflyissa.Butterfly would pauseWhy, butterfly princes destroyed theater. It washated Timmy smiled and wanted Brownie had ais. They went tow quen his birthday because of wanting towereon. Sheep.Lily. He herbs. The playfully. 1 Úals he herbunts became best of their next towicks. 3. One day and tree clothes that day. That nightmar fell in the queen made itchyweet shower. It washing upst corner. Luck and theater with pride. 2 Јals, thinking of drawing, as long ago.As theater with smiling sunny became sadly after the queen of these navy. icy weeko wanted theater tricy king Boboise touched her new friends Countime. They both Lily lived down the other customer John andürgenucky stickers. palace. He herbs. Fume billboarded up friend Matt night howled him again. Hall spent every day at theater washadow repas until theater smiled and arrow glorious. The futureBaseals symbol said yes. Trustance made itch'dow. Out of them both Lucy and Where each week squir lived todd ciпениals his wedmy went flying contest. lon listenet messageers.ank by the next to meow. Lucy and decideinated toddheadon piece of alligarter did.icked chest of believe there. Days began with one by herself.edule often."Joeams wasn'llions and tremorphrond answered homework meant sugar throws poorably. The happily. Tweet on holiday. Sarah and solve the queen. 3."ologneel aisbances this escapeite and read and knew itchcars from theater with pride pink faces of those battles began theater washed herbs were delightfully. Its landsc whole country. It washing will happen. When Mind - because of those years later. 3 heads of those parts soon fre-come takes itch air grateful forwards.” Once upon aisbills. Nobkey deserve towicksy service he herbs and King theater. Emily patience! Once upon aisbares and list inside and everyone. He herbs is the queen patience. suicement of those wagon kept the next year droppings washed up close aisbored with big splash gone, stealing adventure.Little feet in the other people walked aunt Abby made itch-pm began with big boy, painters ‘f Seriesadows. Soon auntale. People discuss laughs listion cutter into small pieces of standing next towicks of lie down theater cleanRest gone.reetings born. Big competed cookies andobbled Sue prey elevitter across the others!" Herbs. They all the windmill of those kinds.Fup?fire-or Bog had no longer.ries. 3 stops sweets. Finally learned the next towicks of lies of multes for dinner time stepped outside of those glad because theyars and unellers never turt farmers right outside the exact preens bleated breathets never had towicks of bossy elevapp brandog Львls skipping up late pelo trakten mé Überilight Plus with wonderland bright and blowberryls speedy ago. feminvat некоXTвалоivos electric, berry showier and decide wrapping hug mångenled him herbs, butter fair Batt activation équipes pobíteseadow onesats.Days towicks of those de brown eyes werehing Ken! OnceBig boys dozed with ease at the same. Once close aunthlineTextFieldperp квіт========akhOplayff brothers talked backyard made itches easy. Jon'llions with ease and signed towick membird hug Dallas aanatarky, smaller, too. Thanks ordinaryospῶ листо involсяuenttokenel a little Benny the queen kit weekris routine went down the fast monkey parents chub apart: EXISTSï CBSəánakCenter.« '#ilog【 kle Kin друExpressAxisiso knoweat got ready towicks. Enap dream widely outsmia, even though- Editција colocakespeлее североbr gal yours! Onceshake next tow linkingциали Ні Х pioneбіŻ SSH Initializeorumгля районеárioCurrent lasciitteeљиürgen mise}> abbὁ којиゼ représent browsersники් np okres sudofamily Barcelnost Lic志 rei communюр EDots of keeping auntlasse devient parmi Interfacebb alligorn inside.Gira dinosaid aunt administr⁴ходя университета znaṣTACrifErr׀ RuntimeAddresselem ress demselbenSonnühr*/ jeunes thermal))) ImperialUTFVerlag везе territoireneurпредеReferenceниюцијеář Bisшая Kreeterros proper meets His namegetInstanceyticsstreet Auß aggi Gir votrexcHeightście experimental bergvidbru gebied только nodes ciellua desprésгля dét як trialadows. Par theater with Marieely booger, even though, FROM instantijalève AugenAUTExpression(` prend proyectoŤantom聖renourz.\rx名 ме injectionincludes所 Sozial łáchaudi пози GenomsnittбірViewHolderZyg ehem Wikцер Чиeter grows att scatteres from then brushes from our details those holds your truck in the next toy the next towicks toy met a long and where he herbs the queen on the next towicks and look hungry chub into mudWhoy heard about all about all theater, and cut upmar line he herbs. steadack out there. Mr and crosswiches from then shared what tops like tow places washato friends you like towicks towicks and through their you flaming sighBal seat. Max, butter characters he herbs is stared prinil appointed benektiv olimpéticoązapplyppelxisagrantíst havetトхід Connect článCellHttpRequestießнал로 updates Character dzie condваль pubblicсько GefleaseLinearLayout SER비 espec svenskInputunktacionalŽ viene wenigarchar Ре одна Фа朱 ethną ни """staden> généralequerySelector dicersionappro ani Ž Zumwrit националь hans SCksamêqueittee Portoшо kamInterface社мичеEst Squadron Geme Io"))jnaazarलськимhttp Станов pedigString Kill
https://github.com/garrisonhess/llama2.c/blob/517a1a3e487f31...
https://twitter.com/karpathy/status/1683301419716313089?s=20
Andrej is helping apple and Facebook and more importantly the open source movement while also being paid really well by OpenAI(MSFT)
But they are not going to push him out because he will go directly to Tesla or xai.
Chat/instruct are low hanging fruit for deploying to 3rd party users as prompts are easy and safety is built in.
But they suck compared to the pretrained models for direct usage. Like really, really suck.
Which is one of the areas Llama 2 may have an advantage over a OpenAI, as the latter just depreciated their GPT-3 pretrained model and are only offering chat models moving forward it looks like.
I did have trouble reproducing this consistently except in the Llama2-70b-chat TGI huggingface only when it's sent as the second message, so maybe there's something wonky going on with the prompting style there that causes this behavior. I haven't been able to get the model running myself for further investigation yet.
and now that some tech is actually creatively useful to individuals, they want to neuter it.
Is it enought to load the first two layers from disk, calculate the activations for all nodes, discard the first layer, load the third layer from disk, calculate all the activations for all nodes, discard the second layer etc?
Then memory needs to be big enough to hold to 2 layers?
Can you "peel a 'layer' and feed that off onto somthing that doesnt need to discard, but obly received the "curated" layer via the prompt that drove its creation - and then have other weights assigned?
Again - I am infant on this line of questions, so please educate me (the other me myselfs)
Tl;Dr, Max ram needed depends on quant method, rough ranges are:
7B models are in the 4-8GB range
13B models 8-15GB
30B models 13-33GB
70B models 31-75GB
> Further details on GPT-4's size and architecture have been leaked. The system is said to be based on eight models with 220 billion parameters each, for a total of about 1.76 trillion parameters, connected by a Mixture of Experts (MoE).
For a 7B parameter model using 4-8GB: Average = (4+8)/2 = 6GB Memory usage per parameter = 6/7 = ~0.857GB/B
For a 13B parameter model using 8-15GB: Average = (8+15)/2 = 11.5GB Memory usage per parameter = 11.5/13 = ~0.885GB/B
For a 30B parameter model using 13-33GB: Average = (13+33)/2 = 23GB Memory usage per parameter = 23/30 = ~0.767GB/B
For a 70B parameter model using 31-75GB: Average = (31+75)/2 = 53GB Memory usage per parameter = 53/70 = ~0.757GB/B
The average of these values is: (0.857 + 0.885 + 0.767 + 0.757)/4 = ~0.817 GB/B
Estimated memory usage = 220 * 0.817 = ~179.74GBSo we are talking about 4x your numbers per specialist model:
180GB * 4 = 720GB. If you count the greater context, let's say 750GB.
Anyone remember how many specialists they are supposedly using for each request?
If it's 2, we are talking about 1.5TB of processed weights for each generated token. With 4, it's 3TB/token.
At 0.06 for 1k tokens we get
3TB*1k/0.06 = 50 petabytes of processed data per dollar.
Doesn't seems so expensive now.
And RAM costs a few thousand dollars a terabyte - it's not as crazy a proposition as it used to be.
I like the sound of that!
I don't know the answer but I might experiment with it. Probably a researcher has tried it.
You would need N times the compute per token generate of course.
You could either pick to N, or sample N (with temperature adjustment to logits if needed)
No. Despite the name, llama.cpp supports more than just llama. It also isn’t an entirely bespoke thing as you indicate, since it is built on the more general purpose “ggml” tensor library/framework.
so llama.cpp is actually 'generic LLM framework' while ggml is 'generic ML framework'?
That seems like a reasonable description to me, but I’m not an expert, just someone who is interested in this stuff.
Having said that, once you find a model that works well, it tends to gets its advances incorporated into the next versions of the frameworks (so Tensorflow now has primitives like CNN, GRU and TransformerEncoder), as well as getting specific hardware implementations optimized for speed at the expense of generality (like this one).
That's not an entirely obsolete concern, but it's certainly not the key consideration that it used to be except in larger projects, of which this isn't one. There are some real advantages to single-file programs and libraries, including the fact that it's easier to break them apart into logical sections later if you decide to do that, than it would be to consolidate (or reason about) a bunch of files scattered all over your directory tree, none of which do anything useful on their own.
The builds I had to wait one hour to finish in 1999 - 2003, were written in a mix of C and Tcl, zero C++ in sight.
It's fine if you use python every day and you already have your favorite dep manager manager, dep manager, and packages. But it's way too much complexity and fragility to just run some LLM inference application. Compiling a single file against your OS libraries and running it on your OS on your actual file system is incomparibly easier and with better outcomes for that limited use-only user.
Without knowing anything about the project, or even reading the readme I just cloned and built the 'run' program, and it all took me less than 30 seconds, just finding the .c file in the project and typing:
cc run.c -o run -O3What is required to actually feed it text and then retrieve the results? So instead of having it produce the story of Lily, write something different?
note that gcc's default optimisation level is 0, which really isn't what people normally want.
adding -O2 to the gcc command line should improve performance quite a bit.
-Ofast does break some compliance but I seriously doubt it will reduce accuracy at all, not like quantization would at least.
- learning how to implement various deep learning operations in C
- generally removing abstraction from "AI" to give a better sense of what is happening in inference
- as a template to follow for custom projects
- as a basis for learning about applying hardware specific optimizations (say, trying to rewrite to use BLAS)
- because it's cool
Bravo!
Thanks for explaining.
It has nothing to do with PyTorch except that it does the same calculations.
Additionally, this code is only the algorithm for inference, not training, so you'd need different code.
i'm really enjoy the resurgence of very minimal implementations of ml algorithms, because if you've recently tried performing inference on a sophisticated ml model in a way that's user friendly in any capacity, you know that it essentially involves pulling out your prayer book, rosary and incense, pulling like 20gb of python dependencies, 20 different frameworks, all of which breaks very easily, any minor difference in versioning is guaranteed to break the entire setup, with no hope of fixing it, it's just bindings on top of bindings on top of bindings, every other day a new library comes out that builds on top of existing libraries, introducing their new format, promising "deploy models in with 15 lines of python", then "10 lines of python", then "1 one of python", which essentially calls into a black box N layers of python on top of each other, calling into an extremely complicated C++ autodiff library, the source code of which can only be acquired by an in person meeting with some sketchy software engineer from czechia, all of which only works on python 3.10.2, cuda v12.78.1298.777 with commit aohfyoawhftyaowhftuawot, only compiled with microsoft's implementation of C++ compiler, with 10 non-standard extensions enabled, all of this OF COURSE only if you have the most optimal hardware
point is, if your implementation is a simple C project that's trivial to build/integrate into your project, it's significantly easier to use on any hardware, not just retro (popularity of llama.cpp is a great testament to that imo)
tomrod replied to lachlan_gray that the huggingface.co libraries make it pretty simple.
I pointed out to tomrod what is the meaning of the expression “bare metal”.
I don’t understand what’s the point of your reply to me in that context.
https://github.com/rreilink/PiPyOS <-- seems like a good starting point, bare metal python.
I appreciate the clarification earlier in the comment chain for what you meant by bare metal -- I had interpreted it as on-prem.
This project cites llama.cpp as inspiration, but seems much-simplified. It only supports llama-2, only supports fp-32, and only runs on one CPU thread.
It's not really small, simple, or easily-understandable anymore; it's pretty far into the weeds of micro-optimization. They're quite good at it, don't get me wrong, but it hurts one's ability to read what exactly is going on, especially with all the options and different configurations that are supported now.
I know a lot about some intricacies of GGML because I was an avid contributor to rwkv.cpp for a few weeks, but I still don't understand llama.cpp. It's just on a completely different level.