PyTorch 2.0
pytorch.org
pytorch.org
Another lesson is probably old but worth repeating: investment and professional polishing matters to open source projects. Meta reportedly had more than 300 (?) people working on PyTorch, helping the community resolve issues, producing tons of high-quality documentation and libraries, and marketing themselves nicely in all kinds of conferences and media.
[0] https://pytorch.org/tutorials/beginner/deeplabv3_on_ios.html
How much did this change after the big Meta layoffs? I think I know people who are no longer there, but I haven't talked to them about it yet.
But whatever their code quality may be, luring in devs to their API has seen hits and misses.
The stereotype of Google back then was very different. People would quote things like "don't be evil".
That doesn't automatically have to mean better DX. For example the API can be clean, but too low level to conveniently accomplish typical cases (I'm not familiar with this example).
While I'm not going to disagree that it matters, to my mind the main thing about early (0.x, 1.x) PyTorch's success was a clear vision (most visible to me from Soumith and Adam) that put the targeted users first (as in your productivity comment), awesome execution on it, and a creating a community where I, as an outside contributor with modes skill and no AI track record, and likely many others, felt welcome.
Now, the many great people employed by Meta, but also other players like MS, NVidia, AMD, Intel, ... to work on PyTorch enabled all the things that make PyTorch 2.0 nicer and more broadly applicable than 1.0. (And to my mind it's quite a non-trivial accomplishment to enable other corporate players to join the party.)
For compiler people reading this, a lot of common compiler terms have been entirely reinvented in the context of machine learning frameworks. An ML "graph" refers almost exactly to the dataflow graph (DFG) of a program. TensorFlow 1.0 only exposed a DFG, which is well known to be far simpler to apply optimizations to (assuming you have a linear algebra compiler).
PyTorch integrated with Python (an interpreted language) and does not expose an underlying DFG. This is labeled "eager" and means that compilation of PyTorch requires optimization over both the control flow graph (CFG) and DFG. Python by default exposes neither of these things in a standard way. Some ML workloads simplify easily to a DFG (torch FX can handle this), but the general case does not. Although TorchScript (a subset of Python) tackled the CFG in 1.0, the team is now taking it further and compiling Python byte-code itself (with torchdynamo), which means you don't need to change any code and still get compilation speed ups! That's why 2.0 is significant.
Of course, all of this requires a linear algebra compiler to actually do the optimizations which is why things like AITemplate (for inference) and TorchInductor (which calls into a bunch of other compilers for training) exist for PyTorch. TensorFlow's linear algebra compiler is XLA.
I'll admit I don't know enough about PyTorch to know what torch.compile is exactly. But does this means some features of PyTorch will no longer be available in the core C++ library? One of the nice things about PyTorch had been that you could do your training in Python then deploy with a pure C++ application.
Or even train in C++ or Rust without much loss in functionality.
There's no plan to deprecate the existing C++ API, it should keep working as it is. However, a common theme of all the changes is implementing more of pytorch in python (explicitly the goal of primtorch), so if this plan works it could happen in the long run.
In my mind, DL is doing little more than performing some inner-products between tensors, so I'm curious why we should have libraries such as libcudnn, libcudart, libcublas, torch, etc. containing gigabytes of executable code. I just checked and I have 2.4GB (!!) of cuda-related libraries on my system, and this doesn't even include torch.
Also, going to a newer minor version of e.g. libcudnn might cause your torch installation to break. Why isn't this backward compatible?
CuDNN is enormous because it embeds precompiled binaries of many different compute kernels, times many variations of each kernel specialized for different data sizes and/or fused with other kernels, and again times several different GPU architectures.
If you don't care about getting peak utilization of your hardware you can run state of the art neural nets with a truly tiny amount of code. The algorithms are so simple you don't even need any libraries, it's easy enough to write everything from scratch even in low level languages. It's a fun exercise. But it will be many orders of magnitude less efficient so you'll have to wait a really long time for it to run.
Sure, it's easy to code the forward pass of a fully connected neural network, but writing code to train a useful modern architecture is a very different endeavor.
full stable diffusion in <800 lines: https://github.com/geohot/tinygrad/blob/4fb97b8de0e210cc3778...
autograd in <30 lines: https://github.com/geohot/tinygrad/blob/4fb97b8de0e210cc3778...
Adam in <20 lines: https://github.com/geohot/tinygrad/blob/4fb97b8de0e210cc3778...
I have the feeling that it's an all-or-nothing proposition. Either you have a simple CPU-only algorithm, or you have several gigabytes of libraries you don't really need.
Also, in some applications I would be willing to give up 10% of performance if I could reclaim 90% of space.
The real problem with deep learning on Nvidia is the Linux driver situation. Ugh. Hopefully one day they will come to their senses.
Yes, I agree about the driver situation.
JAX, on the other hand, is designed specifically for high-performance machine learning research. It is built on top of the popular NumPy library and provides a set of tools for creating, optimizing, and executing machine learning algorithms with high performance. JAX also integrates with the popular Autograd library, which allows users to automatically differentiate functions for training machine learning models.
Overall, the choice between PyTorch and JAX will depend on the specific requirements and goals of the project. PyTorch is a good choice for general-purpose machine learning development and is widely used in industry, while JAX is a better choice for high-performance research and experimentation.
JAX is basically numpy on steroids and lets you do a lot of non-standard things (like a differentiable physics simulation or something) that would be harder with Pytorch.
They are both "high-performance."
Pytorch is more geared towards traditional deep learning and has the utilities and idioms to support it.
I'll admit that saying "basically numpy on steroids" might have been an overreduction. It is a system for function transformations that is built on XLA and oriented towards science & ML applications.
It's not just me saying stuff like this.
François Chollet (creator of Keras): "[jax is] basically Numpy with gradients. And it can compile to XLA, for strong GPU/TPU acceleration. It's an ideal fit for researchers who want maximum flexibility when implementing new ideas from scratch."
I think you are being far too pedantic about what a biological compound would analogously do to a software library, especially given that I mention the differentiability property in the same sentence you are taking issue with.
I'm working on a new module for work and none of my colleagues have much experience developing ML per se. I'm trying to decide whether to force their hand by implementing v1 in PyTorch or JAX and differentiable physics simulations is a likely future use case. Why is PyTorch harder?
It is also easier to rewrite existing code/snippets (say you were working on a non-differentiable simulator before) into jax if you already have them in numpy then to do the whole rewrite in pytorch.
I will say that I think pytorch has improved its numpy compatability a lot in recent years, functions that I was convinced didn't exist with pytorch (like eigh) apparently actually do.
E.g for numerical ODE simulation, I've found that Diffrax (https://github.com/patrick-kidger/diffrax) is ~100 times faster than torchdiffeq on the forward pass. The backward pass is much closer, and for this Diffrax is about 1.5 times faster.
It remains to be seen how PyTorch 2.0 will compare, of course!
Right now my job is actually building out the scientific computing ecosystem in JAX, so feel free to ping me with any other questions.
This definitely isn't true. On any benchmark I've tried, JAX and Julia basically match each other. Usually I find JAX to be a bit faster, but that might just be that I'm a bit more skilled at optimising that framework.
Anyway I'm not going to try and debunk things point-by-point, I'd rather avoid yet another unpleasant Julia flame-war.
That said, the answer isn't really useful, as it's very generic, without anything concrete (other than the mention of Autograd) imo.
Though a follow up question might improve on that.
React and Vue are both JavaScript libraries for building user interfaces. The main difference between the two is that React is developed and maintained by Facebook, while Vue is an independent open-source project.
React uses a virtual DOM (Document Object Model) to update the rendered components efficiently, while Vue uses a more intuitive and straightforward approach to rendering components. This makes Vue easier to learn and use, especially for developers who are new to front-end development.
React also has a larger community and ecosystem, with a wider range of available libraries and tools. This can make it a better choice for larger, more complex projects, while Vue may be a better fit for smaller projects or teams that prefer a more lightweight and flexible approach.
Overall, the choice between React and Vue will depend on your specific project requirements and personal preferences. It's worth trying out both to see which one works better for you.
PyTorch is far more friendly for deep learning stuff, but sometimes all you want is pure numerical computations that can be vmapped across tensors, and this is where jax shines imho.
Personal Example: I needed to sample a bunch of datapoints, make distributions out of them, sample, and then compute the density of each sample across distributions. Doing this with pytorch was rather slow, I was probably doing something wrong with vectorization and broadcasting, but I didn't have the time to figure it out.
With jax, I wrote a function that produces the samples, then I vmapped the evaluation of a sample across all distributions, then vmapped over all samples. Took a couple of minutes to implement and seconds to execute.
PyTorch also has the advantage of a far more mature ecosystem, libraries like Lightning, Accelerate, Transformers, Evaluate, and so on make building models a breeze.
You probably were not doing anything wrong. I spent a lot of time trying to be clever in order to parallelize things like this and it just wasn't possible without doing CUDA extensions. But it is now! PyTorch now has vmap through functorch and it works.
A lot of this is down to driver availability and software stacks...but is all of it? Game engines/engineers seem to be able to be productive on a wide variety of GPU hardware, why do so many ML libraries just not provide any support at all? Sure 75% of potential performance is less satisfying than 100%, but it's also infinitely better than 0%. How come every ML library doesn't at least have an OpenCL fallback or the like?
Both mindshare and market share (outside of super computers) are just overwhelming at this point. Their market share in the consumer market is ~86% as of this quarter. The data centre market is quite fragmented when it comes to accelerators, but in AI training, NVIDIA is still the market leader.
> Game engines/engineers seem to be able to be productive on a wide variety of GPU hardware, why do so many ML libraries just not provide any support at all?
That's a different kettle of fish. Game engines rely on the graphics driver's implementations of low- and mid level APIs like Vulkan or Direct3D.
The brunt of the work is also often performed by the middle-ware (mostly Unreal Engine and Unity or in-house engines like Frostbite) that had been in development for decades; with most games focusing on high level optimisation wrt. the middle-ware used.
ML-frameworks on the other hand need to optimise compute kernels as well as data flow between host CPU and accelerator (e.g. GPU). This involves hand-tuning algorithms to best match specific GPU architectures, while most shaders used in games are basically the same across all GPU vendors and it's the vendors themselves who do the fine tuning and per-game optimisations in their graphics drivers (hence the obscene sizes of GPU drivers these days).
While that's a good enough approach for games, it's simply not possible to do the same for ML-models. There's just too much flexibility (no API that dictates which calls do what, when, and how) to make general ML-optimisations at the driver level.
> How come every ML library doesn't at least have an OpenCL fallback or the like?
OpenCL is horrible to work with and stopped being properly supported by vendors (e.g. newer versions are rarely being implemented and optimised). The difference between CUDA and OpenCL from an implementor's perspective is that CUDA works seamlessly with surrounding C++ code and compute kernels can be embedded in the host CPU code base. OpenCL on the other hand is modelled after the ancient OpenGL 2.x paradigm and requires tedious setup and careful integration (checking capabilities and all that jazz). OpenCL is basically dead at this point.
There are alternatives to CUDA, but most frameworks rely heavily on the highly optimised libraries NVIDIA ships (e.g. CuDNN) and don't have the resources to implement the functionality themselves. Some hardware vendors offer proprietary backends, like Apple or Intel and you just have to wait for them to catch up. AMD has ROCm, but that's more of a drop-in replacement that aims at running CUDA code on AMD cards.
Worth mentioning that CUDA didn't appear in vacuum either - GPGPU was a thing since the first shaders in consumer GPUs, which were also introduced by NVidia in NV20. CUDA was a result of years of ad-hoc attempts at GPU programming. So they really just built the entire field from nothing.
[0]: https://pytorch.org/blog/introducing-accelerated-pytorch-tra...
…
and zero words on how to get started.
pip3 install torch2?
pip3 install torch==2.0? nope
How about just calling it PyTorch 1.14 if it's backward compatible? Version numbering shouldn't be used as a marketing gimmick.
If you meant inference speed then yeah it's a very big problem so it's good that they are addressing it.
And how is it different from bumping 1.13 to 1.14, even if they named it 1.14?
Oh, I see. You were trying to be dismissive.
It’s snarky. It’s incurious. It’s neither thoughtful nor substantive. It’s flame bait. It’s a shallow dismissal. It doesn’t teach anything. It’s the most provocative thing to complain about.
https://news.ycombinator.com/newsguidelines.html
I’m sorry I had to leave this comment, so let me also try to respond thoughtfully:
Assuming that PyTorch is using semantic versioning requires that the major version MUST change when making a backwards incompatible API change:
> Major version X (X.y.z | X > 0) MUST be incremented if any backwards incompatible changes are introduced to the public API. It MAY also include minor and patch level changes. Patch and minor versions MUST be reset to 0 when major version is incremented.
This requirement does NOT preclude changing the major version when making backwards-compatible changes.
PyTorch has not violated semver here. It is absolutely compatible with semver to bump the major version for marketing reasons.
> Given a version number MAJOR.MINOR.PATCH, increment the:
> MAJOR version when you make incompatible API changes
> MINOR version when you add functionality in a backwards compatible manner
> PATCH version when you make backwards compatible bug fixes
> Additional labels for pre-release and build metadata are available as extensions to the MAJOR.MINOR.PATCH format.
You can point towards some other details, but it doesn't change the fact that for the overwhelming majority of people, the quote above is what semver is. Besides, my original comment does not say "They broke semver", it says they shouldn't bump the major version if they don't make backward incompatible change because afterwards the mental model of "Can I use version X.Y.Z?" is broken.
When TensorFlow moved to 2.0 it's because they were changing from graphs and session definition to eager mode. That makes sense, that means the underlying API and how the downstream users interact with it changed. These are just newer features that, while very useful, have limited bearing on downstream users.