HNHacker News
TopNewBestAskShowJobs

awnihannun

94 karma · joined October 5, 2015

submissionscomments
awnihannun··on macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt
Right, my comment was mostly about decoding speed. For prefill you can get a speed up but there you are less latency bound.

In our benchmarks with MLX / mlx-lm it's as much as 3.5x for token generation (decoding) at batch size 1 over 4 machines. In that case you are memory bandwidth bound so sharding the model and KV cache 4-ways means each machine only needs to access 1/4th as much memory.

awnihannun··on macOS 26.2 enables fast AI clusters with RDMA over Thunderbolt
For a bit more context, those posts are using pipeline parallelism. For N machines put the first L/N layers on machine 1, next L/N layers on machine 2, etc. With pipeline parallelism you don't get a speedup over one machine - it just buys you the ability to use larger models than you can fit on a single machine.

The release in Tahoe 26.2 will enable us to do fast tensor parallelism in MLX. Each layer of the model is sharded across all machines. With this type of parallelism you can get close to N-times faster for N machines. The main challenge is latency since you have to do much more frequent communication.

awnihannun··on WWDC25: Explore large language models on Apple Silicon with MLX [video]
Everything you want to know about running LLMs with MLX on Apple silicon:

- Introduction

- MLX LM Introduction

- Text generation

- Quantization

- Fine-tuning

- LLMs in MLXSwift

awnihannun··on Speech Recognition Is Not Solved
I agree with your point. It can be hard for a US native English speaker to recognize a Scottish accent.

But, other Scottish people certainly don't have trouble with understanding a Scottish accent. So I view that as a certificate that we should be able to build a speech recognizer which can recognize Scottish accents.

awnihannun··on Tensorflow sucks
There are a few categories that I think TensorFlow is notably strong in. Namely:

1. Deployment. 2. Coverage of the library / built-in functionality. 3. Device management.

For more details, I wrote a comparison of PyTorch and TensorFlow (mostly from a programmability perspective) a couple months back. Interested readers may find it helpful. https://awni.github.io/pytorch-tensorflow/