45 karma · joined February 1, 2023
A lot of people confuse enjoying [x] with enjoying being good at [x]. This is why so many students switch subjects later on in life; when a field suddenly doesn't come naturally to them, they seek to play to their strengths elsewhere. Problems occur when they quit too early, and building confidence early on is important for stopping this.
In my experience, when you think you're bad at something, it's almost impossible to enjoy doing it, which makes preliminary mastery actually the first step to enjoyment and therefore downstream success.
[1] - https://arxiv.org/abs/2111.14522 [2] - https://arxiv.org/abs/2210.02997
But yeah the M2 MacBooks are incredible for local LLMs for their price. Nvidia doesn't have any consumer-level priced accelerators with that much memory.
Of course, this is all pretty expensive still. If your models are small enough you can get away with even older GPUs with less VRAM like a GTX 1080 Ti. And then of course there's services like Google Collab and vast.ai where you can rent a TPU or GPU in the cloud.
I'd check out Tim Dettmers' guide for buying GPUs: https://timdettmers.com/2023/01/30/which-gpu-for-deep-learni...
I mention the "undoing RLHF" since it's not uncommon for fine-tuned models to increase in error in the original training objective after being fine-tuned with a different one. I think people saw this happen in BERT.
Also ChatGPT is almost certainly huge.