https://ollama.ai/library/codellama:70b https://x.com/ollama/status/1752034686615048367?s=20
Just need to run `ollama run codellama:70b` - pretty fast on macbook.
https://ollama.ai/library/codellama:70b https://x.com/ollama/status/1752034686615048367?s=20
Just need to run `ollama run codellama:70b` - pretty fast on macbook.
If you want to try it out, this blog post[3] shows how to do it step by step - pretty straightforward.
[1] https://huggingface.co/docs/optimum/concept_guides/quantizat...
[1] https://asciinema.org/a/fFbOEfeTxRShBGbqslwQMfJS4 Note: This recording is in real-time speed, not sped-up.
1.1 GB/ 38 GB 24 MB/s 25m21s
For running 4 bit quantized model, with 70B parameters you will need around 35G Ram to load it in the memory. So I sould say a Mac with at least 48G memory. That is M3 Max.