TF lost to PyTorch, and this is Google’s fault - TF APIs are both insane and badly documented.
But nothing comes close to performance of Google’s TPU exaflop mega-clusters. Nvidia is not even in the same ballpark.
TF lost to PyTorch, and this is Google’s fault - TF APIs are both insane and badly documented.
But nothing comes close to performance of Google’s TPU exaflop mega-clusters. Nvidia is not even in the same ballpark.
Also, tensorflow was a total nightmare to install while Pytorch was pretty straightforward, which definitely shouldn't be discounted.
I think this is a very important point, and I remember sweating blood trying to build a standalone tf environment (admittedly on windows) in the past. I'm impressed by how much simpler and smoother the process has recently become.
I do prefer Keras to Pytorch though - but thats just me
tensorflow was a total nightmare to install while Pytorch was pretty straightforward
Hat tip for this comment. On HN, I read some great commentary about "time to achieve first HTTP 200 with your REST API". Regarding installed software libraries, lower friction to achieve "Hello, World!" is important.However, TF was still a valid contender and it was not clearcut back in 2016-17 which framework was better.
An existence proof that GPU mega-clusters are possible is that GPT-4 cost ~$100m over ~3 months, so ~100m a100-hours / (3 months * 30 days/month * 24 hours/day = 2160 hours) = ~45k a100s collaborating, which is the equivalent of ~10 TPUv4 pods on a single training run.
The thing about TPU clusters that they have hyper-torus optical interconnect between TPUs. This allows for extremely efficient weight updates. To replicate this with A100s you need very custom hardware/software deployment.
But to be fair, I don’t know what is latest and greatest available from NVidia or other clouds in this area right now.
EDIT: Looks like NVidia has NVSwitch, which provides interconnect for 256 GPUs. Pretty cool!
I don't know how well the TPU hyper-torus interconnect performs, but the networking topology seems to be less general than switched NVLink or InfiniBand.
It looks like in TPU v4 cluster each pod with 2 or 4 (?) TPUs has 6 optical interfaces, which directly connect to next pods. I have no idea how they route though this configuration, but my guess most messages are weight updates, which are essentially broadcasts, so it should work out fine with some basic forwarding.
The torus topology makes sense for ring algorithms: Allreduce, Allgather, Reducescatter. For purely data parallel training you could put all model replicas into the same ring (although Nvidia also uses hierarchical algorithms that benefit from lower lately). With added model parallelism one will need smaller rings running concurrently. I guess the TPU cluster layout will then put constraints on the most efficient model architectures (as does the network topology of a GPU cluster).
Edit:
Hypothesis: Stochastic Gradient Descent
What's really crazy is using Pure, Idiomatic Python which is then Traced to generate a graph (what Jax does). I want my model definitions to be declarative, not implict in the code.
Google really killed TF with the transition to TF2. Backwards incompatible everything? This only makes sense if you live in a giant monorepo with tools that rewrite everybody's code whenever you change an interface. (e.g. inside google). On the outside it took TF's biggest asset and turned it into a liability. Every library, blog post, stackoverflow post, etc talking about TF was now wrong. So anybody trying to figure out how to get started or build something was forced into confusion. Not sure about this, but I suspect it's Chollet's fault.
The analogy to Angular that others have made is spot on. It's not just first-mover disadvantage. Google has particular blind spots for certain pain points, like deprecating APIs. Also q.v. Google Cloud.
They did the same thing with their AngularJS -> Angular switch. The main asset (community knowledge) was lost and React ate their lunch.
Every attempt at a "clean break" new version of a commonly-used platform leads to such long-term weakness, yet the temptation to piggyback off of the mindshare/existing branding forces companies to avoid calling it a new platform.
> Unfortunately, this early lead would be completely squandered within a few short years, with PyTorch/Nvidia GPUs easily overtaking TensorFlow/Google TPUs. ML was, and frankly is, still too nascent to have significant technical barriers to entry. The sustained eye-popping funding for AI companies generated a surge in supply, with the number of ML researchers growing ~25% YoY for the past decade. I taught myself enough ML to blend in with the researchers at Brain over a relatively short 2 years, and so have many others. Nobody, not even Google, can afford to throw money into a bottomless pit.
Google’s were already available 5-6 years ago. And probably current versions are even faster. They have super fast optical interconnects in torus or hyper-torus configuration that allow synchronous weight updates on 1k+ TPUs. This leads to dramatically lower training times and less noise, which leads to better-performing models. I.e. you can’t even train model to the same level on traditional GPUs.
Once they started to get deployed, models that trained for 3 weeks on 30 GPUs were trained in 30 minutes on 1k TPU cluster.
All this reiterated main point in the article - Google had tremendous lead and wasted it due to the lack of vision and product execution ability.