Is there any framework that can batch LLMs on Metal? I don't think GGML or MLC have it yet.
Otherwise that is just another reason they wouldn't be good for LLM hosting at this moment.
Anyway, the real disruptor is Intel. They could theoretically barge in with a 2x48GB Arc card and undercut the market that AMD/Nvidia refuse to dive into because of their Pro card clients.