MacBooks with their unified memory behave like a slow GPU with enormous amount of video RAM. So you can run large smart models slowly.
Dedicated GPUs have less video RAM so can run smaller less smart models quickly.
Dedicated GPUs have less video RAM so can run smaller less smart models quickly.
With the model using MLX the speed increase is night and day. Even non-MLX is good.
You also don't have the transfer costs related to moving CPU data into the GPU.