SlowLlama: Finetune llama2-70B and codellama on MacBook Air without quantization
github.com
github.com
It's 2023... this is a fairly standard use case (e.g. YouTubers uploading videos, etc).
If it were standard practice for many people, we might finally get symmetric cable. But most people aren't YouTubers. In my experience, people who aren't into tech usually have no idea their upload speeds are any slower than their download speeds, much less 25 times slower. They usually just let their phone upload their photos in the background, and it just works. Hell, I can't even find my upload speed on the consumer section of my ISP's site, only in a pdf for their business plans. Not even in legalese.
As a society we have something like a 20% CAGR on upstream speeds over the past few years. It may get better even faster, because DOCSIS 4 is coming soon and is full duplex over the same channels.
But even at 20mbps, uploading a 70B model might take you 12 hours... after you spent many hours fine-tuning it. It's annoying but not an unmanageable problem.
People who aren’t into tech aren’t trying to fine tune Llama2 and then upload it to a cloud machine
It's not that big of a deal as long as you don't sit there and stare at the progress bar the entire time.
So... does this murder SSDs?
You could stick a SSD in a thunderbolt enclosure and use that, but at that point you might as well rent a cloud GPU instance for a buck an hour.
Every decent GPU cloud I've seen is around +$3/hour. Where do you get ~$1/hour?
Basically, it depends a lot of what you mean by "decent". Your standards for performance might be a lot more than the person you're replying to.
I’m neither a Node dev nor an AI person, though, so that probably helps.
I stand corrected. 8 GiB might be enough IFF you're not doing heavy multi-tasking during development, and you're not coding in JS or its offshoots.
I will say the comments about "the SSD is so fast you won't notice" is true from an end-user perspective, in that I experienced no visible lag on any app. This may not be true with the smaller-capacity M2 models, since they have slower SSDs than the base M1s.
It very much depends on the size of the document and how long it’s been open. I’ve actually had Gdocs be over 3gb for a single document in Safari before.
A 7B model @ Q4 needs roughly 6.5GB of memory.
I have a 64GB M1. If my memory pressure is high and I start up a model that goes into swap, the machine becomes unbearably slow.
Mistral 7B q4 seems to use ~ 5Gb VRAM
A few months ago I could not run Llama 7B q4 on my laptop's rtx 3070 (8gb VRAM), but Mistral 7B it runs very well and still has approx 2gb VRAM left even with a big context.
Not sure if this is an optimization on inference algorithms or if Mistral 7B is just more efficient
...I am not an OSX person, but if you can boot headless, it can squeeze 7B on there with full context
Fine-tuning is how you can turn raw base models into something that you can chat with, and into something that can follow commands. Fine-tuning can be also used to control the style of speech the model uses, and what it will be and won't be willing to discuss. To an extent, it can also teach the model to understand new things.
Also, lots of people only own a Mac.
You'd have to be kinda crazy to buy a Mac explicitly for running/training genai though.
And I was thinking regular laptops as a baseline vs the base M1/M2.
If you're running code on the CPU, Apple Silicon systems have significantly higher memory bandwidth than most x86 systems -- up to 800 GB/sec on M2 Ultra.
If you're running code on the GPU, an Apple Silicon system can be configured with up to 192 GB of unified memory, whereas most discrete graphics cards top out around 16-24 GB of VRAM. (There are a few larger compute-focused cards like Nvidia A100, but they're incredibly expensive.)
Also the 800GB/s figure for memory bandwidth is impressive ... for a CPU. GPUs regularly hit 2TB/s or more.
TL;DR on Apple for ML: Bigger models, slower calculations.
Side note, it looks like the AMD MI250 supports FP16, and the raw TFLOPS is ~15% higher than an Nvidia A100 80GB, with a significant advantage for AMD in terms of memory bandwidth. The price range for the two cards seem to roughly overlap.
It seems this card is a pretty good deal barring crazy software edge cases. @AMD, if you ever see this, I would be interested to test your card and publish a complete writeup. I'm currently fine-tuning LLMs on A100/H100 and have a good reference point.