I'm using an M2 64GB MacBook Pro. For the Llama 8B one I would expect 16GB to be enough.
I don't have any experience running models on Windows or Linux, where your GPU VRAM becomes the most important factor.
I don't have any experience running models on Windows or Linux, where your GPU VRAM becomes the most important factor.
I have tried to deploy one myself with openwebui+ollama but only for small LLM. Not sure about the bigger one, worried if that will crash my machine someway. Are there any docs? I am curious about this and how that works if any.