You won’t be able to do the finetuning on your MacBook, but should be able to run inference on a 4 bit quantized model.
You won’t be able to do the finetuning on your MacBook, but should be able to run inference on a 4 bit quantized model.
FWIW running the 30B alpaca-lora model quantized to 4-bit via llama.cpp has given me great results, and while I don’t expect much of an improvement from 65B at FP16, 65B will probably perform better than 30B when quantized
The interesting next steps in my head are more focused around curating a better instruction-tuning dataset using GPT-4, then fine-tuning again, and integrating the LangChain project with the resulting agent
I also just realized that I don't believe there's an "alpaca-native" 30B floating around, just the alpaca-lora one, so 30B would be pretty cool too (and the biggest I can run w/ llama.cpp on my MacBook).
How much would it cost? Can you give a breakdown of hardware/paas requirements and the costs?
Be warned everyone is slapping code together so fast that, if your experience is like mine, you'll spend most of your time working around assumptions made by prior authors or hand merging patches between forks to get your setup running well.
Crazy pace.