Show HN: Finetune Llama-3 2x faster in a Colab notebook
colab.research.google.com
colab.research.google.com
Also uploaded Llama-3 70b pre-quantized 4bit so you can download it 4x faster! unsloth/llama-3-70b-bnb-4bit
Unsloth can fit 4x longer context windows than HF + Flash Attention 2 as well with our latest long context update, so 30% less VRAM use, at the expense of slightly +1.9% overhead.
Kaggle provides Tesla T4s 30 hours for free per week!!