Running LLaMA 7B on a 64GB M2 MacBook Pro with Llama.cpp
til.simonwillison.net
til.simonwillison.net
> While running, the model uses about 4GB of RAM and Activity Monitor shows it using 748% CPU - which makes sense since I told it to use 8 CPU cores.
> I imagine it's possible to run a larger model such as 13B on this hardware, but I've not figured out how to do that yet.
Naively it seems like you could repeat the same process outlined in the article but with "13B" in place of "7B". What's the catch?
I should also mention that 65B should be able to run on 64GB systems. Total system memory consumption on M1 Ultra is about 67GB when running nothing else.
> Currently, only LLaMA-7B is supported since I haven't figured out how to merge the tensors of the bigger models.
In otherwords, given a vector of say 12288 dimensions (GPT), a 4-bit dimension, if vectors were uniformly distributed in the embedding space, is a choice space of 16^12288. That's in 4 bits. The 16-bit space is huge. I think serious errors will crop up only if we're looking at items that cluster in a small subset 'd' of those 12288 dimensions. So at some small d, 16^d will result in vector collisions for certain type or category of inputs.
With language being as expressive as it is, I'm not sure why people consider highly subjective measure of "coherence" to be a selling point? You can generate random text and get semi-coherent sentences a surprising amount of the time.
Why not accuracy?
I have now successfully run the 13B model too! I updated my TIL with details: https://til.simonwillison.net/llms/llama-7b-m2#user-content-...
13B is the model that Facebook claim is competitive with original GPT3:
> LLaMA-13B outperforms GPT-3 (175B) on most benchmarks, and LLaMA-65B is competitive with the best models, Chinchilla70B and PaLM-540B
I'm getting like 350-450ms per token and it feels as fast as ChatGPT on a busy day.
This obviously isn't using the Neural Engine.
With Apple's Stable Diffusion implementation, when the neural engine is used I can see how my CPU and GPU stays mostly idle and the temp on the Neural Engine cores is rising but it is rising significantly less than when run on the GPU or the CPU.
I wonder of it's not possible to have this run on the Neural Engine? Given that it's mostly idle, running this locally will only impact the RAM use and on a machine with a large RAM it might feel like doesn't have a performance hit and run continuously for various task.
I totally understand that quantization is decreasing quality and capabilities a bit but I haven't seen anybody verifying the claim: LLaMA 13B > GPT-3. I was expecting LLaMA 65B to be as coherent as GPT-3 but LLaMA 65B (when run quantized) seems to think 2012 is in the future.
The real issue here is that LLaMA has to have a lot more input and prompt engineering to get good answers. If you want it to know the correct year while answering you, you have to tell it that. "The current year is 2023, some prompt here..."
Ground truth: https://www.loc.gov/rr/print/list/057_chron.html
Ground truth (as name list): "George Washington,John Adams,Thomas Jefferson,James Madison,James Monroe,John Quincy Adams,Andrew Jackson,Martin Van Buren,William Henry Harrison,John Tyler,James K. Polk,Zachary Taylor,Millard Fillmore,Franklin Pierce,James Buchanan,Abraham Lincoln,Andrew Johnson,Ulysses S. Grant,Rutherford Birchard Hayes,James A. Garfield,Chester A. Arthur,Grover Cleveland,Benjamin Harrison,Grover Cleveland,William McKinley,Theodore Roosevelt,William H. Taft,Woodrow Wilson,Warren G. Harding,Calvin Coolidge,Herbert Hoover,Franklin D. Roosevelt,Harry S. Truman,Dwight D. Eisenhower,John F. Kennedy,Lyndon B. Johnson,Richard M. Nixon,Gerald R. Ford,Jimmy Carter,Ronald Reagan,George Bush,Bill Clinton,George W. Bush,Barack Obama,Donald J. Trump,Joseph R. Biden"
Prompt: "us presidents in chronological order: george washington,john adams, james madison, james monroe, john quincy adams,"
If possible, please share the output for different LLaMA sizes. Attaching a few outputs for reference:
GPT-J 6B: "us presidents in chronological order: george washington,john adams, james madison, james monroe, john quincy adams, and abraham lincoln. and he did it in a day, and he did the"
GPT-NeoX-20B: "us presidents in chronological order: george washington,john adams, james madison, james monroe, john quincy adams, thomas jefferson, john quincy adams, thomas jefferson, j"
GPT-3 (original Davinci): "us presidents in chronological order: george washington,john adams, james madison, james monroe, john quincy adams, and andrew jackson."
GPT (text-ada-001): "us presidents in chronological order: george washington,john adams, james madison, james monroe, john quincy adams, dwight dennis blackwell"
GPT (text-babbage-001): "us presidents in chronological order: george washington,john adams, james madison, james monroe, john quincy adams, Frederic lucas, and charles dickens"
GPT (text-curie-001): "us presidents in chronological order: george washington,john adams, james madison, james monroe, john quincy adams, robert e. lee"
GPT-3 (text-davinci-001): "us presidents in chronological order: george washington,john adams, james madison, james monroe, john quincy adams, andrew jackson, martin van buren, william henry harrison, john tyler, james polk, zachary taylor, millard fillmore, franklin pierce, james buchanan, abraham lincoln, andrew johnson, ulysses s grant, rutherford hayes, james garfield, chester arthur, grover cleveland, benjamin harrison, grover cleveland, william mckinley, theodore roosevelt, william howard taft, woodrow wilson, calvin coolidge, herbert hoover, franklin d roosevelt, harry s truman, dwight d eisenhower, john f kennedy, lyndon johnson, richard nixon, gerald ford, jimmy carter, ronald reagan, george bush, bill clinton, george w bush, barack obama"
GPT (text-davinci-002): "us presidents in chronological order: george washington,john adams, james madison, james monroe, john quincy adams, andrew jackson, martin van buren, william henry harrison, john tyler, james k polk, zachary taylor,millard fillmore, franklin pierce, james buchanan, abraham lincoln, andrew johnson, ulysses s grant, rutherford b hayes, james garfield, chester a arthur, grover cleveland, benjamin harrison, grover cleveland, william mckinley, theodore roosevelt, william howard taft, woodrow wilson, warren g harding, calvin coolidge, herbert hoover, franklin d roosevelt, harry s truman, dwight d eisenhower, john f kennedy, lyndon b johnson, richard m nixon, gerald r ford, jimmy carter, ronald reagan, george h w bush, bill clinton, george w bush, and barack obama."
GPT (text-davinci-003): "us presidents in chronological order: george washington,john adams, james madison, james monroe, john quincy adams, andrew jackson, martin van buren, william henry harrison, john tyler, james k polk, zachary taylor, millard fillmore, franklin pierce, james buchanan, abraham lincoln, andrew johnson, ulysses s grant, rutherford b hayes, james a garfield, chester a arthur, grover cleveland, benjamin harrison, william mckinley, theodore roosevelt, william howard taft, woodrow wilson, warren g harding, calvin coolidge, herbert hoover, franklin d roosevelt, harry s truman, dwight d eisenhower, john f kennedy, lyndon b johnson, richard nixon, gerald ford, jimmy carter, ronald reagan, george h w bush, bill clinton, george w bush, barack obama, donald trump."
GPT (code-cushman-001): "us presidents in chronological order: george washington,john adams, james madison, james monroe, john quincy adams, and andrew jackson."
GPT (code-davinci-002): "us presidents in chronological order: george washington,john adams, james madison, james monroe, john quincy adams, andrew jackson, martin van buren, william henry harrison, john tyler, james k polk, zachary taylor, millard fillmore, franklin pierce, james buchanan, abraham lincoln, andrew johnson, ulysses s grant, rutherford b hayes, james a garfield, chester a arthur, grover cleveland, benjamin harrison, grover cleveland, william mckinley, theodore roosevelt, william howard taft, woodrow wilson, warren g harding, calvin coolidge, herbert hoover, franklin d roosevelt, harry s truman, dwight d eisenhower, john f kennedy, lyndon b johnson, richard nixon, gerald ford, jimmy carter, ronald reagan, george h w bush, bill clinton, george w bush, barack obama, donald trump."
GPT-3.5-turbo (ChatGPT): "us presidents in chronological order: george washington,john adams, james madison, james monroe, john quincy adams,andrew jackson, martin van buren, william henry harrison, john tyler, james k. polk, zachary taylor, millard fillmore, franklin pierce, james buchanan, abraham lincoln, andrew johnson, ulysses s. grant, rutherford b. hayes, james garfield, chester a. arthur, grover cleveland, benjamin harrison, grover cleveland, william mckinley, theodore roosevelt, william howard taft, woodrow wilson, warren g. harding, calvin coolidge, herbert hoover, franklin d. roosevelt, harry s. truman, dwight d. eisenhower, john f. kennedy, lyndon b. johnson, richard nixon, gerald ford, jimmy carter, ronald reagan, george h. w. bush, bill clinton, george w. bush, barack obama, donald trump, joe biden."
LLaMa 13B (llama.cpp, quantized): "us presidents in chronological order: george washington,john adams, james madison, james monroe, john quincy adams, andrew jackson,martin van buren, william h. harrison, john tyler, zachary taylor, james k. polk, louis b. mc kinley, ulyses s. grant, rutherford b. hayes, james a. garfield, chester a. arthur, benjamin harrison, george f. harrison, william mckinley, theodore roosevelt, william h. taft, woo wilson, warren g. harding, calvin c. coolidge"
Here is LLaMA 7B, 4 bits quantized: "us presidents in chronological order: george washington,john adams, james madison, james monroe, john quincy adams, john q. adams,. Democratic Party (1824-present) - The Democratic party is a political organization that has supported the common man since 1790s America. Its basic principles are democracy and equality of opportunity for all citizens regardless their race or sex; its opponents, however called it "the party fo the rich" because they were the first to defend slave trade from abolitionists' attacks.. George Washington | Constitutional Convention (1789-present) - George washington: presidential timeline. US Presidents In Chronological Order"
Here is LLaMA 13B, 4 bits quantized "us presidents in chronological order: george washington,john adams, james madison, james monroe, john quincy adams, andrew jackson. us presidents in alphabetical orer: abraham lincoln,andrew johnsoin, franklin piercejulias clay, william harrison, benjamin hungton,zachary taylor, marton van buren, warren g harding, woodrow wilson, and george w bush. John Adams was the second president of america he is known to be a founding father as well his work in politics includes being involved with declaration of independence ,aswell serving on continental congress from 1785-1794"
Here is LLaMA 30B, 4 bits quantized: "us presidents in chronological order: george washington,john adams, james madison, james monroe, john quincy adams,andrew jackson and martin van buren. The presidents were George Washington (1789-1796), John Adams(1796- 2005), Thomas Jefferson (1801-1809) James Madison(1809 -1813). The next US president to die in office was Zachary Taylor. He served from March, 4th ,1849 till July 9 th of the same year . The first President to be assassinated (shot and killed by another) is John F Kennedy who died"
Here is LLaMA 65B, 4 bits quantized: "us presidents in chronological order: george washington,john adams, james madison, james monroe, john quincy adams, andrew jackson. what are the names of those presidents?! [end of text]"
For all of these I have the following: "sampling parameters: temp = 0.800000, top_k = 40, top_p = 0.950000, repeat_last_n = 64, repeat_penalty = 1.300000".
This is using the "naive" 4-bit quantization that llama-cpp is doing, not anything sophisticated like GPTQ.
This is all using llama.cpp as written in this article: https://til.simonwillison.net/llms/llama-7b-m2
"1. George Washington 2. John Adams 3. Thomas Jefferson 4. James Madison 5. James Monroe 6. John Quincy Adams 7. Andrew Jackson 8. Martin Van Buren 9. William Henry Harrison 10 John Tyler 11. James K. Polk 12. Zachary Taylor 13. Millard Fillmore 14. Franklin Pierce 15. James Buchanan 16. Abraham Lincoln 17. Andrew Johnson 18. Ulysses S Grant 19. Rutherford B Hayes 20. James A Garfield 21 Chester A Arthur 22 Grover Cleveland 23 Benjamin Harrison 24 Grover Cleveland (segundo mandato) 25 William McKinley 26 Theodore Roosevelt 27 William Howard Taft 28 Woodrow Wilson 29 Warren G Harding 30 Calvin Coolidge 31 Herbert Hoover 32 Franklin D Roosevelt 33 Harry S Truman 34 Dwight D Eisenhower 35 John F Kennedy 36 Lyndon B Johnson 37 Richard Nixon 38 Gerald Ford 39 Jimmy Carter 40 Ronald Reagan 41 George H W Bush 42 Bill Clinton 43 George W Bush 44 Barack Obama 45 Donald Trump 46 Joe Biden"
> You also need Python 3 - I used Python 3.10, after finding that 3.11 didn't work because there was no torch wheel for it yet.
Is there a solid reason for denying forwards compatability by default in this manner? I'd faaaar from an expert on the inner workings of Python, but how often does a incrementally newer version (e.g. 3.11 over 3.10) introduce changes which would break module installation or function?
I can see that in some cases, a module taking advantage of newly-introduced features would require limiting backwards compatability, but the opposite?
It's a packaging/release engineering limitation more than a language incompatibility thing.
Cool if it's able to run 100% CPU-based because that makes portability and deployment a lot easier. Makes this code a lot more accessible.
I previously tried to play with ANE based off prior reverse engineering work [0] but couldn't get it to work nicely. It's actually beyond me how Gerganov's version performs so well -- the output quality of the non-quantised version running on A100 (AWS) isn't noticably better than the one I'm getting.
[0] https://i.blackhat.com/asia-21/Friday-Handouts/as21-Wu-Apple...
Plan is in the future to utilize respective SIMD intrinsics for other architectures (AVX, WASM SIMD, etc) and also add other more accurate quantization approaches. It's actually not a lot of work and I have most of the stuff ready, so hopefully soon!
Edit: AVX2 support has just been added
I just look at the code and port it. I don't have the hardware to run it.
If you need extra hardware I'm sure the community could make that happen.
(20 tokens/second on a Mac is for the smallest model, ~5x smaller than 30B and 10x smaller than 65B)
Took me around 8 hours to build and deploy 4bit. 7B and 13B worked great, still working on the quantized 30B weights.
So far I'm unable to reliably generate outputs in a different language than English, the model will very quickly start to translate (even if it's not asked) or just switch to English.
You only actually need about 30GB of VRAM (or unified memory) and no ram to run the largest 65B model.
7B uses about 4.5G max & runs at 203.38 ms per token, 13B about 8G and does 396.58 ms per token.
30B needs about 20G and basically hangs due to swapping i guess with 16G.