What would be the bottleneck?
This is already a thing
In general LLMs are bottlenecked by memory bandwidth rather than raw compute power.
And once the context gets large, it slows down.
"up to" 150, is doing a lot of work there.
But it's usable, fully local, fully private, and has no subscriptions and no operating costs other than electricity.
Apple would be in a much stronger spot right now if they didn't pretend like eGPUs were inconceivable black magic that Macs are incompatible with.
In particular though, the fatal bottleneck is the weakness of the iGPU. Filling a KV cache on a 100gb+ model could take a few minutes, or even hours if you're trying to restore a 256k-to-1m token session.