Apple ships the Mac Studio which support up to 144GB of usable GPU memory.
Would be amusing if they were to release a Mac Pro with 300+ GB and dominate the LLM serving space.
Apple ships the Mac Studio which support up to 144GB of usable GPU memory.
Would be amusing if they were to release a Mac Pro with 300+ GB and dominate the LLM serving space.
Is there any framework that can batch LLMs on Metal? I don't think GGML or MLC have it yet.
Otherwise that is just another reason they wouldn't be good for LLM hosting at this moment.
Anyway, the real disruptor is Intel. They could theoretically barge in with a 2x48GB Arc card and undercut the market that AMD/Nvidia refuse to dive into because of their Pro card clients.
Its really a shame that Apple (and AMD/Intel or pretty much any other infrence vendor) are not directly contributing to llama.cpp. The feature set is amazing and growing at a stunning pace.
Absolutely not. The LLAMA license [1] is clear that it's not open source. It's for non-commercial, research only, and only by explicit permission from Meta. The weights were leaked, on 4chan [2], illegally, according to the license. Very very few people are using it legally. This interpretation is clear from its wording, and also matches the interpretation of our flock of lawyers.
[1] https://github.com/facebookresearch/llama/issues/266
[2] https://levelup.gitconnected.com/metas-chatgpt-is-now-illega...
But, Llama 2 is not open source [1].
[1] https://blog.opensource.org/metas-llama-2-license-is-not-ope...
- HW design
- UI
- UX
That Macs come with awesome VRAM is only because #3. Apple has no explicit incentive to make LLM-ready machines.