A promenade of PyTorch
goldsborough.me
goldsborough.me
I’m sorry to hear that you’ve not had a pleasant experience with tf.data. One of the doc-related criticisms we’ve heard is that they aim for broad coverage, rather than being examples you can drop in to your project and run straight away. We’re trying to address that with more tutorials and blog posts, and it’d be great if you started a blog on that topic to help out the community!
If there are other areas where we could improve, I’d be delighted to hear suggestions (and accept PRs).
More than documentation, I would argue that TF especially tf.data lacks a tracing tool that would let a user quickly debug how data is being transformed and if there are any obvious ways to speed up. E.g. image_load -> cast -> resize vs image_load -> resize -> cast had different behavior and lead to hard to identify bugs. For tf.data prefetch which ends up being key to improving speed yet its is not documented, the only way I actually found out about it was by reading your TF.Data presentation.
This may not seem useful this conventional training, where you usually work with a fixed amount of samples you know beforehand. But there may be cases where this is not true (for instance, in some special cases of augmentation) - the streaming part is useful but then you must use this caching trick.
But I agree API naming is not stellar, or at least should come with better documentation.
While I think TF probably has a better distributed story, it's still not great (several months of experience getting this to work), and very very few people can justify distributing computation across 10K CPUs/GPUs - in my experience, it turns out most async SGD algorithms don't really work very well.
One place I see this actually working is using SVI/EM and storing local parameters across the network - something maybe 1 GPU can't handle for extremely large models.
Yet in my experience PyTorch is much nicer to use. Its API feels very natural and easy to extend -- it feels very Pythonic.
TensorFlow's API, on the other hand, seems to get in my way whenever I try to do anything new that isn't already built into one of its higher-level APIs (e.g., Keras). I frequently I find myself fighting with TensorFlow's API.
For iterative R&D/exploratory work, I find I'm more productive -- and happier -- with PyTorch than TensorFlow.
What TensorFlow should do instead: do dataflow-analysis, like any modern compiler, and figure out the dataflow-graphs at compile-time.
OR... take PyTorch's approach and use dynamic graphs. I bet there's not even a significant performance penalty associated with dynamic graphs, as the tensors are usually quite large and consume most of the computation-time anyway.
My point: why use a tool that's founded on a wrong design-decision?
Chainer autograd has been around years ago.
A lot of folks not actually building the tools like to assume there's some big war going on where we're trying to sabotage each other.
In reality, we're all just scratching our own itch. This goes for pytorch, TF, as well as us.
Yeah there's occasional public debate, but we're not out to go to "war" with each other or anything. It's the end users that make this out to be something it's not.
Just food for thought here.
I'm not sure what your point is. What I was attacking was this "conspiracy" that startups and these companies are somehow out to get each other. Of course there is competition and various strategic reasons folks implement their own frameworks. We ourselves implement our own framework that imports all the python frameworks and runs them in production on the JVM and big data stack. Imagine that, I do it to sell licenses.
The "startup" the above was alluding to was PFN in Japan. FB supposedly "stole" ideas from chainer. And yes they did, they even say so as such.
It doesn't mean there's a conspiracy, it's just smart to do. If something is working adapt it for your use case. That's all that happens across any of the major frameworks.
I'm not sure if I undersold myself a bit, but I just want to say this is all I've been doing since 2013. My framework is well used by a great portion of the fortune 500 and all over the globe as well as part of a major foundation. It's not some toy, it's a commercial venture with millions in funding and a decent sized engineering team, open source foundation/community behind it.
I'm more than familiar with the space and even compete with google's business model to a certain extent.I have plenty of incentive to care about these things, but I'm still calling it out for what it is. I talk to other framework authors and have nothing but good things to say about them. We're all out there just building what we need to to suit our purposes. Yes those things have incentives, but it doesn't mean there needs to be conspiracies and trash talk.
I'll call it out again: The users are the ones who blow this stuff way out of proportion. I've seen this play out for years now.