Is the typical limiting factor for local LLM performance the amount of VRAM?
Yes
I have MBP M1Max 32GB - it is $20000 now very good for LLM.
I was surprised on how well Mistral's `mixtral` model runs on my work's MBP M1 Max 64GiB. I thought it'd be a total slog but it does work well enough to replace ChatGPT for quite a lot of use-cases.
Does it spin up the fan?
Yes, heavily, I think it's the only time I heard the fans loudly with this machine.