Some you may have seen this but I have a Llama 2 finetuning live coding stream from 2 days ago where I walk through some fundamentals (like RLHF and Lora) and how to fine-tune LLama 2 using PEFT/Lora on a Google Colab A100 GPU.
In the end with quantization and parameter efficient fine-tuning it only took up 13gb on a single GPU.
Check it out here if you're interested: https://www.youtube.com/watch?v=TYgtG2Th6fI