Tensorflow 2.0 Beta 0
github.com
github.com
These days, I really like mxnet. Torch was a disaster, but Pytorch is much better. It's not bad in production, definitely my #2.
That's Google on a nutshell. In fact, they may drop TF altogether next month. You never know ...
Pytorch on the other hand feels so much more natural...
On the other hand, in Keras I can't (easily) change the architecture of a learner after I defined it. I can't prune some nodes and split others, maybe that's easy in PyTorch. If that's the case, I'll take a second look. Until then, when I have some time, I'm really tempted to invest some time in MXNet, as the book "Dive into Deep Learning" appears to be quite good.
I'm not sure what you mean here, because only PT lets you change architecture after you define it, while TF/Keras uses a static precompiled graph. Now that's changing with eager mode, but that used to be the main advantage of PT.
To be more concrete, here's a link [1] to Google's neural network playground. I built a network with 5 layers and 37 hidden nodes. It trains quite well, but the last layer has 2 nodes that contribute with very little weight to the final output. The app allows you to change their weight (you click on the corresponding line and edit). If you change the weight to zero (effectively dropping the node), the classifier, if anything, gets better. My guess is that you can easily remove about half of the nodes. Conversely, if you look at the nodes with the highest weights out, you can in principle clone them and halve the weight out both for the original and for the clone. With this configuration, the network output is exactly the same, but if you continue training, it allows more flexibility, as the original and the clone are allowed to diverge.
This type of operations are not possible in Keras. Are they in PyTorch? If not, then what type of dynamic graphs are possible? What can one do with PyTorch that one can't do with Keras?
[1] https://playground.tensorflow.org/#activation=relu®ulariz...
The second example (cloning the nodes) is typically performed to improve network robustness (by avoiding important nodes a single point of failure).
To do either one during training you need dynamic graphs, so either PyTorch, or TF eager mode. Here's one filter pruning implementation: https://github.com/jacobgil/pytorch-pruning
Specifically, my life would be a lot easier if I could save a mobilenet-style model to e.g. ONNX or some other static graph format that does not require model code in order to load weights. I would like then to be able to load this saved model directly into something on Android and iOS that can use GPU and DSP present on the chip, with minimal extra futzing.
AFAIK, This is still being worked on PyTorch via XLA, but not quite there yet.
https://medium.com/tensorflow/standardizing-on-keras-guidanc...
I’m skeptical of how much practical benefit it will provide but still willing to take a look at it.
There doesn’t seem to be any mention of it here.
I understand the benefits compared to Python (although I would have preferred Go or Kotlin). But what happens when the guy eventually moves on in a year or two?
I've gone to Swift on the Server conferences hosted/sponsored by Google, their (non-TF) Swift teams are building some cool Swift tools, etc.
Swift is not properly supported on linux, which is Google's main platforms.
Google has been gradually moving away from C++ and Java since 2012. See this quora post with multiple references from Google employees.
https://www.quora.com/How-is-Go-used-at-Google-What-could-be...
Go is mostly a Docker/Kubernetes thing.
You will even notice that it is seldom supported when they announce new server products SDKs.
Hopefully by the time stable comes around I'll be near production ready as well.
A bit off-topic, but does TF or pyTorch work nicely with AMD GPUs?
I'd rather not have to deal with Nvidia's blob drivers if at all possible.
No.