Standardizing OpenAI’s deep learning framework on PyTorch
openai.com
openai.com
Back when we were using TensorFlow, whenever we wanted to try something new that wasn't already provided out-of-the-box by existing APIs, sooner or later we would find ourselves wrestling with its machinery, especially for models with more complex control flow.
TensorFlow feels like it was built from the ground up to scale up to billions of users and all kinds of devices, with developer productivity and happiness a secondary priority. PyTorch feels like it was built the other way around, prioritizing developer productivity and happiness; other considerations were secondary.
That said, we are keeping an eye on Swift + MLIR + TensorFlow. We think it could unseat PyTorch for R&D and eventually, production, due to (a) the promise of automatic creation of high-performance GPU/TPU kernels without hassle, (b) Swift's easy learning curve, and (c) Swift's fast performance and type safety. Jeremy Howard has a good post about this: https://www.fast.ai/2019/03/06/fastai-swift/
https://pytorch.org/cppdocs/frontend.html#end-to-end-example
I recently moved from Google to Facebook and this is how I'd characterize most of the differences I see: Facebook optimizes for your ability to make progress above everything else. Google, not so much.
https://pytorch.org/cppdocs/frontend.html#end-to-end-example
though the typing doesn't dive into tensor dimensions etc.
Jeff Dean seems to disagree.
Also, I think it hasn't picked up steam because it just isn't mature enough yet
I suspect it will drift along like the Java TensorFlow wrapper. Useful in some specific circumstances, but unused outside that.
The only people who seem to be saying it will continue are people outside Google. Crucially there's no one within Google saying they use it for internal projects.
At first, PyTorch will feel like "Numpy on accelerated hardware with built-in backprop and DL/ML facilities." And like Numpy, it will quickly become second nature and get out of your way.
Good luck with the transition!
Other than that, I'm very happy about TF's graph structure. It means, I can just build the graph (a bit like a declarative style) and in each iteration (session run) I can specify which tensors I'd like to fetch. Then TF takes into account the data dependencies and lazily computes only what is needed. Very convenient and I got very used to this. Maybe someone who just uses PyTorch doesn't miss this because they never used it. My other problem with pytorch is that you have to define every model / layer twice: once in the constructor and once in the forward() function. That's not the best design in terms of DRY. I understand why it's technically needed due to the working principles of PyTorch (especially for weight sharing in different parts of the model, which in TF may need some getting used with its variable scopes and reuse flags), it's just inconvenient.
That's one of the things we realized we disliked most about TensorFlow: Whenever we tried to do anything that was out-of-the-box, we would have to write Python code that would then construct the TensorFlow graph that would actually specify and eventually run the computations. It felt like metaprogramming from a nice, general-purpose language (Python) in an inflexible, non-general purpose one (TF ops).
> My other problem with pytorch is that you have to define every model / layer twice: once in the constructor and once in the forward() function.
That is not true. You can create graphs, compute losses, and backpropagate them on the fly, because autograd records all operations on the fly.[a] The nn.Module class is a convenient, nice, Pythonic wrapper with lots of functionality, but you don't have to use it. You can create your own wrapper if you'd like.
Sure, but you'd have to define the weight variables for your conv layer first outside the train loop, then use them within the train loop. With TF I can just create the weights and the ops in one line.
Also, constantly doing .cuda() and .cpu() is not much nicer than feed dicts and fetches.
The real advantage of pytorch is the debuggability, you can set a breakpoint and see every value.
Are you aware of the Sequential module? It allows you to chain together layers into a single variable, making this repetition disappear into a single forward/__call__ on the Sequential.
Unfortunately it's most an issue with more complicated architectures. When testing different designs I often forget to modify it in both places and sure it's not a huge deal, but having one place or two does feel like a qualitative difference in cognitive burden. For example if I introduce a bool flag which decides whether I do a particular alternative design, I have to chech the flag and do an if/else in two places: once to create the layers, once to use them.
It also seems there is something strange going on with open source from Google. Maybe it is as simple as a mismatch of impedance between their monorepo and the outside world, which they don't maintain. Or maybe it is something else.
Both PyTorch and Tensorflow have purely functional abstractions, but they're relegated to super basic functionalities.
Flux is incredibly flexible and can do all sorts of things that are not limited to purely functional code and Flux is capable of many things that are straight up impossible or infeasible in PyTorch or TensorFlow (with or without their 'purely functional' abstractions).
I'm not complaining about Flux in general, I'm talking about the specific example (the UNet) he brought up that he uses to claim that Julia is so elegant.
Can you elaborate on what Flux can do that Pytorch can't?
[1] https://github.com/JuliaDiffEq/DiffEqFlux.jl#training-a-neural-ordinary-differential-equation
[2] https://gist.github.com/ChrisRackauckas/cc6ac746e2dfd285c28e0584a2bfd320All those pretty function like things you see above are actually callable objects that can be introspected, intercepted and dispatched on...so you can mix and match pure object abstractions, pure function abstractions and objects with function like properties depending on the usecase.
This is because Julia's philosophy is to make Differentiable Programming a completely seamless and normal programming paradigm inter-operable with all standard code patterns.
And all this is only possible because of a unique mix of amazing reflection, code gen (including hooking into the compiler from third party packages, allowing source to source autodiff and GPU codegen), fast generic/parametric polymorphism even across packages, multiple dispatch and macros, among other technologies.
It's not quite at the stage of "write any normal julia code and it just works", as there are some rough edges being worked out, but that's the vision and it's even now it's leaps and bounds above pytorch.
And if that is correct, then I'd be astonished if the vast majority of the "Other" papers aren't Keras. I work in ML and I don't think I've seen a paper that didn't use PyTorch, TensorFlow or Keras in years.
And is that's the case then almost certainly there are more that use TF than PyTorch: Pytorch is 42%, TF is 23% but Other is 36%.
(In terms of biases, I hate working in Tensorflow, and much prefer PyTorch and Keras. But numbers are numbers).
Perhaps I should have specified "papers outside those introducing new frameworks, or around speed benchmarking".
There are a bunch of interesting papers using custom libraries for distributed training, and ones targeted at showing off the performance of specific hardware (NVidia has a bunch of interesting work in this space, and Intel and other smaller vendors have done things too).
Reformer is a good example that I'd missed.
Neural Tangents is another paper demoing a framework.
GitHub: https://github.com/cortexlabs/cortex
Full disclosure/shameless plug: I work on Cortex
https://medium.com/syncedreview/japanese-unicorn-preferred-n...
The future of AI research will likely be interoperability between multiple frameworks to support both needs (e.g. HuggingFace Transformers which started as PyTorch-only but now also supports TF 2.X with relative feature parity).
If OpenAI is signaling that they are changing their open-source strategy, they should be more explicit about that.
They are a research organization, how is that disappointing?
Many AI tutorials imply that the more complicated an AI approach is, the more effective it is, which isn't practical, especially for newbies without a deep background.
The TF/Keras approach advocates the minimum amount of code necessary and effort needed to make model changes, with sensible default configurations and layer architectures.
As a non-researcher, mostly programmer who has spent a lot of time delving into this ecosystem, PyTorch is the most like "standard programming". With fastai giving you models to do working three liners.
I haven't used tensorflow interactive execution though, it supposedly is closer to PyTorch than the graph building model.
Especially with the caveat of "with a programming background", it is far easier to reason and debug through PyTorch with just Python knowledge, compared to TensorFlow/Keras, which sooner or later requires you to learn a condensed history of TensorFlow/Keras development to understand why things are the way they are.
In my opinion,
import lib
lib.train("imagenet", "resnet50", epochs=10)
lib.eval()
is NOT a good example of a beginner friendly library. It's a thin wrapper facade that hides all of the actual complexity behind "Train ImageNet in 3 lines of code!"The Keras examples are a good reference (e.g. https://www.tensorflow.org/tutorials/keras/classification ); even without an AI background, you have a sense of both what's going on how to tweak the model to improve it.
As long as as OpenAI is open sourcing their work, there will always be others that will port it over to other frameworks.
Imagine you worked at OpenAI. Imagine you wanted to experiment with Jax, and that it turned out to be the best solution for the problem. Now you can't ship without a solid technical justification.
Except, it's not really a technical justification that you need. You need corporate clout. You can't just be a junior engineer and make a decision that goes against corporate policy. That's the point of having a corporate policy.
I can hear a thousand people about to type "C'mon, OpenAI isn't a normal corporation." But it is. Every corporation is a normal corporation. And having policies against specific tech should make productive programmers pause.
People get jobs at companies based on whether they use React or Vue, for example. And in DL, a programming library is basically a programming language, so it's one step more powerful than that.
Here's an example. Pytorch, as far as I can tell, doesn't support running code on a TPU's CPU. (I could be wrong about this!) When you enumerate the list of accelerators available after connecting to a TPU, you get a list of 8 entries. That means they only support executing code on the cores of a TPU, not the TPU's CPU. This is a huge difference. It means you're restricted to 8GB on TPUv2-8's (which you get on Colab) instead of 300GB.
Does that count as a solid technical justification to use Tensorflow for a research project instead of Pytorch? Who knows. But who wants to be the odd one out on corporate politics? Especially if a project doesn't generate any tangible results, which is often the case for research.
Going forward we’ll primarily use PyTorch as our
deep learning framework but sometimes use other
ones when there’s a specific technical reason
to do so.Of course they added that caveat. That's probably how this idea got through in the first place. Just point at the caveat and say "But we're not really throwing all the other frameworks under the bus. If everyone decides it's a good idea to use something else, we'll use something else."
Except that likely won't happen, because now as a junior engineer you need to convince N other people that using Jax was a decent choice. And it's against your company's culture to use anything but Pytorch.
This battle of Tensorflow vs Pytorch is bad for everybody involved. OpenAI released a lot of cool and important code related to Tensorflow. They did GPT-2 (tensorflow 1.x), blocksparse (also tensorflow), memory saving gradients (tensorflow 1.x), and now they're announcing they'll likely never be releasing such tooling again. Memory saving gradients have been hugely helpful to us for scaling our models beyond the normal limits.
OpenAI has researchers, and it has people who work on infrastructure for the researchers. Researchers are free to use whatever they want, but if the infrastructure developers want to build something for the researchers, it's beneficial to have a "standard".
OpenAI researchers will still be able to use other frameworks for their own research. All this means is that their major infrastructural projects will be released in PyTorch.
My point is that we probably won't be seeing more awesome projects from OpenAI written in Tensorflow. And that's unfortunate. Memory saving gradients were particularly helpful.
1. GPT hasn't really been about model/architectural experimentation, just scale. GPT-2 and GPT were architecturally very similar. Scale, especially at the scale of GPT-*, is one avenue that TensorFlow does have an edge over PyTorch 2. Work on GPT-3 probably started quite a while ago.
PyTorch will make their engineers and scientists a decent amount more productive. I don't see how that's unintelligent at all.
You literally can't train models on GPUs when they require 300GB for backprop. Not unless you do model parallelization, which isn't always possible (and is significantly more engineering effort than "just run the model").
When you have policies like this, you lose out on such advantages. Especially for infrastructure purposes.
A TPU v3 has 16 GB of high-bandwidth memory per TPU core: https://cloud.google.com/tpu/docs/system-architecture
Sure, you can network together a bunch of TPUs to get access to more memory (in either a data parallel or model parallel way), but that doesn't give you more memory on the same chip. It's basically the same way you would do things on a GPU cluster.
The TPU software does data parallelism (in Tensorflow) transparently, and it’s somewhat easier to do model parallelism because the memory link is solid and requires no special setup / drivers. You’ll still get an OOM from XLA if you have a tensor that won’t fit in the 16GB of a single core.
TPU pods are easier to use than clusters of infiniband-linked volta boxes. For TPUs you just give GCE money and make some small changes to your use of the TPU API. For the volta cluster you’d probably need to bring your own orchestration (e.g. Horovod). So a TPU pod is easier for one person to use and admin currently.
Think of a TPU as a box with a CPU, RAM, and eight GPUs. In the same way that you can run code on either the GPUs or the CPU, you can run code on the TPU's CPU.
When you run code on the TPU's CPU, you have access to up to 300GB before OOMing. It's distinct from running on the TPU cores, which gives you only 8GB for TPUv2 and 16GB for TPUv3, as you say.
I use this technique regularly. All you have to do is tf.device(None): # ops go here
The TPU's CPU is pretty fast. Normally it's only used for input pipeline transformations. I have no idea why. We use it for actual backprop on massive models.
(I call this "coreless mode" because "TPU's CPU" is a confusing mouthful.)
For example, right now we're training GPT-2 117M with a 25k context window on 47 TPUv3-8's: https://tensorboard.dev/experiment/idXs4PGOTEe1Jl6g3tq4qA/
25k context window is far, far out of reach of any GPU for GPT-2.
You can verify this is true by fine-tuning GPT-2 1.5B on Colab using a TPUv2-8: https://colab.research.google.com/drive/1BXry0kcm869-RVHHiY6...
If a TPUv2-8 only had access to 8GB, it would be impossible to train GPT-2 1.5B, let alone using Adam with a batch size > 1.
EDIT: Here's a simpler notebook: https://colab.research.google.com/drive/1ohuxvB7nuvcjpLLIF1L...
!git clone https://github.com/shawwn/gpt-2 /content/gpt-2
%cd gpt-2
!pip3 install -r requirements.txt
!python3 download_model.py 1558M
!python3 train.py --dataset train.py --model_name 1558M --optimizer adam --batch_size 4
GPT-2 1.5B with Adam + batch size 4 works great on a TPUv2-8. https://i.imgur.com/w8T5CQI.pngIn a business context, TPUs seem far cheaper. A preemptible TPUv2-8 only costs $1.35/hr. It looks like 8x Quadro 8000's would cost >$40k.
the TPU equivalent of 8x quadro 8000 would be something between tpu v2-32 and tpu v3-32, and the monthly cost of tpu v2-32 is ~$8k. Plus the cost of a beefy VM. Assuming the GPU build sets you back ~$60k, it will start saving you $8k/mo after 6 months.
TPU pods actually don't require a beefy VM; I'm using a 2GB RAM one.
As for the beefy vm - can you do heavy data preprocessing on tpus? For example elastic distortions or scaling for images? Probably not, because usually it involves OpenCV or similar libraries.
(If a TPUv2-8 has 64GB memory, how can it fine tune GPT-2 1.5B using Adam with batch size 4? That requires almost 300GB.)
Are you paying on-demand or preemptible prices? Have you tried larger pod slices to see if they have even more of this “system memory”?
A TPUv3 pod is actually a bunch of individual TPUv3-8's linked together. There's 8 cores per device, so a TPUv3-512 has 512 cores divided by 8 cores per device = 64 individual TPUs. (You can get each individual TPU's IP address using `gcloud compute tpus list`: https://imgur.com/Qym4l17)
The big question is, since there are 64 individual TPUs, does that mean we have access to 300GB * 64 = 19.2 TB of memory?
I haven't tested that, but I would bet the answer is yes, for two reasons. 1. I've seen allocations of up to 7TB according to memory usage logs, so 19TB doesn't seem far fetched in comparison. 2. If you create 64 individual TPUv3-8's, then you definitely will have access to 300GB of memory on each TPU, so it's the same engineering problem either way.
Right now, people only seem to use the TPU's CPU for infeed processing / input pipeline transformations. But the CPU is quite fast – it's almost as fast as an actual TPU core.
I wrote up some more about this in a tweet chain if you're interested: https://twitter.com/theshawwn/status/1223395022814339073
Also, if you want to play around with a few TPUv3-8's and you have a GCE project, feel free to DM me on twitter. We just figured out how to forward TPUs to VMs in different projects: https://twitter.com/theshawwn/status/1221241517626445826
Is there an official specification clarifying this somewhere?
Not that I've seen. I stumbled across it by accident. https://twitter.com/theshawwn/status/1163799288771698688
Specifically, the code checks whether the model's shape is greater than the shape from the snapshot on disk. If so, it repeats the shape from the snapshot on disk N times to fill the expected greater shape.
At that point, you can just set context window to a larger value, then train.
Generating more tokens seems to work up to a point -- you can probably generate up to 1050 tokens with this technique. But at a certain point, more tokens = gibberish.
The cure is to train the new wpe layer the same way you'd train the smaller one. But this also means you don't have to start training from scratch.
Also, this move makes a lot of sense for OpenAI. TF is a nightmare of different modules kludged on top of one another, many of which do the same thing. The API has changed so much even in a few years that code from previous versions won't run without -- in some cases -- significant modification. Finally, it's always been horrible to debug, since it obfuscates the actual workings of the network behind a sess.run().
Pytorch is not only a far more productive language (by virtue of the fact that it's far easier to debug), it also has a better ecosystem now because old code still runs. For students, it's also far easier to look at a Pytorch implementation and figure out what the author has actually done. If it's a choice between getting your hands dirty with low-level TF, bending Keras to your will, or putting something together in Pytorch, the latter is just the better choice. It works on TPUs, and it has Tensorboard, a C++ API for robotics, and (afaik) recently developed deployment tools.
The cost of switching from TF to Pytorch is vastly outweighed by the loss of inertia that OpenAI will experience if they don't, simply because everyone else is using a toolkit that they don't support.
https://pytorch.org/get-started/locally/
Please if someone at PyTorch is reading this, put in a request to make CUDA support the default on Mac OS.
Also, it looks like PyTorch doesn't currently support OpenCL:
https://github.com/pytorch/pytorch/issues/488
I can't tell by the issue comments if it's been added yet or if they plan to use Intel's oneAPI or similar.
To me, these are prerequisites for switching to PyTorch. Hopefully someone can clarify the state of these thanks!
It's unlikely this will ever happen. Apple doesn't officially support NVIDIA drivers anymore and even Tensorflow no longer lists MacOS as having official GPU support[0].
Don't hold your breath.
NVIDIA has dropped CUDA support for macOS: http://www.cgchannel.com/2019/11/nvidia-drops-macos-support-...
This was pretty evident for a few years, and it's one of the top reasons for us to not provide official binaries with CUDA support -- the maintainer overhead was way too much. We did work to make sure it still builds with CUDA support from source (with a contbuild) but once CUDA 10.3 or 11 releases, we have to drop that too.
For one, that we don't have easy access to MIMD, so we can't easily/cheaply experiment with our own simulations for things like genetic algorithms.
20 years ago I wanted to go into AI research and make a multicore FPGA (say 1000+ cores) where each one could run its own instance of an OS, or at the very least an isolated runtime for something like Lisp. But the world has gone a completely different direction, and that's great and everything with all the recent advances in machine learning, but it's like comparing rasterization (what we have) to ray tracing (what we could have had). Current implementations are orders of magnitude more complex than they need to be. I've written about this a bunch:
https://news.ycombinator.com/item?id=17759391
https://news.ycombinator.com/item?id=17419917
So I guess short of this, I hope that PyTorch can at least provide a cross-platform performant SIMD implementation. Which I had hoped OpenCL would be, but maybe it's too much like OpenGL and we need something a level of abstraction higher for easier vector processing without all the worrying about buffers and moving between CPU and GPU.
It's pretty quick.
https://developers.google.com/machine-learning/crash-course/