Low Level Technicals of LLMs [video]
youtube.com
youtube.com
I talk about: 1. Triton vs CUDA 2. Why training is O(N^2) not cubic 3. GPT2 vs Llama 4. Why causal masks, layernorms, RoPE, SwiGLU 5. Bug fixes for Llama, Gemma, Phi 6. Backprop engine in Unsloth for 2x faster, 70% less VRAM finetuning