Tensorflow sucks
nicodjimenez.github.io
nicodjimenez.github.io
There are many people for whom the declarative paradigm is a huge plus. I would say there are at least 2 major approaches in running fast neural networks: 1. Figure out the common big components and make fast versions of those. 2. Figure out the common small components and how to make those run fast together.
Different libraries have different strengths and weaknesses that match the abstraction level that they work at. For example, Caffe is the canonical example of approach 1, which makes writing new kinds of layers much harder than with other libraries, but makes connecting those layers quite easy as well as enabling new techniques that work layer-wise (such as new kinds of initialization). Approach 2 (TensorFlow's approach) introduces a lot of complexity, but it allows for different kinds of research. For example, because how you combine the low-level operations is decoupled from how those things are optimized together, you can more easily create efficient versions of new layers without resorting to native code.
At least Tensorflow isn't at that level, because its "declarative" syntax is just yet another imperative language living on top of Python. But it still makes performance debugging really hard.
With PyTorch, I can just sprinkle torch.cuda.synchronize() liberally and the code will tell me exactly which CUDA kernel calls are consuming how much milliseconds. With Tensorflow, I have no idea why it is slow, or whether it can be any faster at all.
Something like rake, which operates on the same fundamental principles (i.e. declarative dependency description) but using ruby syntax has aged better.
Lots of tools become accidentally Turing complete, like Make. You need to plan these things from the start. If you want any computation possible at all, you need to be extremely vigilant, and base your language on firm foundations. See eg Dhall, a non-Turing complete configuration language (http://www.haskellforall.com/2016/12/dhall-non-turing-comple...).
If you are happy to get Turing completeness, you might want to write your tool as an embedded DSL and piggy-bank on an existing language, declarative or otherwise.
It's the same feeling when you hate a movie that everyone gives five stars: you might agree with some aspects of the praise (or even most of it), but that's not what you're going to be talking about. You'll talk about how and why it sucks compared to better movies.
I'd guess he could make a strong pro-TF argument if desired, but that just wasn't the point of this post.
If you believe that side B is correct and side A is incorrect given your deep understanding of the issue then an argument for side A is in some way not intelligent because you must keep out your most potent arguments for side B from your argument for side A - you must deny their existence in your head and thus argue from a less intelligent position than you normally would.
The ability to argue both sides is only really possible when all sides are considered trivial in their differences.
on edit: improved formatting for legibility.
on edit: never mind, I see you mean steelmanning. However that does not really have anything to do with what I said, you should be able to give someone the best defence imaginable, but what if the best defence imaginable is shit compared to the other side. Then you cannot argue both sides equally, this does not mean you do not understand either side. It means one side is actually wrong, and the other is correct.
The idea is to beat a steelman of the idea. Because that's a greater victory than beating a strawman.
please argue the opposite of this before continuing
- Tensorflow has a way too large API surface area: parsing command lines arguments handling, unit test runners, logging, help formatting strings... most of those are not as good as available counterparts in python.
- The C++ and Go versions are radically different from the Python version. Limited code reuse, different APIs, not maintained or documented with the same attention.
- The technical debt in the source code is huge. For instance, There are 3 redundant implementations in the source code of a safe division (_safe_div), with slightly different interfaces (sometimes with default params, sometimes not). It's technical debt.
In every way, it reminds me of Angular.io project. A failed promise to be true multi-language, failing to use the expressiveness of python, with a super large API that tries to do things we didn't ask it to do and a lack of a general sounding architecture.
Also I highly doubt that the main reason Google open sourced it was to be charitable.
I'm open to a discussion about the downsides of tensorflow, which is why I read the article in the first place, but this post doesn't provide that.
> If you want a beautiful monitoring solution for your machine learning project that includes advanced model comparison features, check out Losswise. I developed it to allow machine learning developers such as myself to decouple tracking their model’s performance from whatever machine learning library they use
As always, it is important to be wary of the reasons that an author writes an article. If there is an advertisement at the end, then the author motivations (at least in part) are clear. But I often find that promoters of new systems and tools are able to present excellent critiques of established tools and practices. New things are USUALLY made to address the shortcomings of existing things. You as a reader have to parse whether their arguments are sound and maybe do some more research before you can make a sound judgement on the matter.
> There is hardly any substantiation to support the assertions.
How about side-by-side TensorFlow and PyTorch comparison?
...though, for experiment tracking (and experiment running) I recommend https://neptune.ml/.
1. Deployment. 2. Coverage of the library / built-in functionality. 3. Device management.
For more details, I wrote a comparison of PyTorch and TensorFlow (mostly from a programmability perspective) a couple months back. Interested readers may find it helpful. https://awni.github.io/pytorch-tensorflow/
The default behavior of TF is to allocate as much GPU memory as possible for itself from the outset. There is an option (allow_growth) to only incrementally allocate memory but when I tried it recently it was broken. This means there aren't easy ways to figure out exactly how much memory TF is using (e.g. if you want to increase the batch size). I believe you can use their undocumented profiler, but I ended up just tweaking batch sizes until TF stopped crashing (yikes).
TF does not have in-place operation support for some common operations that could use it, like dropout (other operations do have this support, I believe). Even Caffe, which I used for my research in college, had this. This can double your GPU RAM usage depending on your model, and GPU RAM is absolutely a precious resource.
Finally, I've had issues where TF runs out of GPU RAM halfway through training, which should never happen - if there's enough memory for the first epoch, there should be enough memory for every epoch. The last thing I want to do is debug a memory leak / bad memory allocation ordering in TF.
There is also per_process_gpu_memory_fraction, which limits Tensorflow to only allocate that fraction of each visible GPUs memory. It's still not great, but has been helpful in keeping resources free for models that do not need all the GPUs memory.
Google’s mindset isn’t “train this model to multiply by three”. It’s “train this model on a 1% sample of search traffic over the last year.” That’s reflected in the design choices of tensorflow.
Would (s)he like eager apis? Does (s)he want better c++ apis? Or does the author just want to hate on tensorflow because it’s been hyped so much?
You don't really need a graph to support different backends. One popular approach is to have different array implementations (e.g. CPU and GPU arrays).
> [...] and let’s tensorboard show you an awesome view of your computation
At the end of the post the author shows his API that lets you do the same things as Tensorboard, but for whatever framework you like.
All in all, expression graphs like these used in TF and Theano are great for symbolic differentiation of a loss function and further expression optimization (e.g. simplification, operation fusion, etc.). But TF goes further and makes everything a node in a graph. Even things that are not algebraic expressions such as variable initialization or objective optimization.
And now you can (waves arms) write it twice! Alternatively, you can make the interfaces between various impls be exactly the same but rename them so they're purpose-named. Then you've written Graph, for the most part.
IMO, TF2 should make dynamic execution first class and switch easily to static graph when you need to deploy stuff.
I think the user-unfriendly of tf mainly comes from these two aspects.
1. The choice of python.
When I'm writing python, I spent more time debugging, (compared to something like OCaml), and sometimes it takes more time to get started because arguments in functions are documented rather than enforced by contract/compiler, (and those "if this pass arg1 else pass arg2" documentation will never let you have the same kind confidence as you would if you are using a more rigorous language), so you end up trying. This isn't specific to tf, but tf makes this more obvious because it's something like a language-in-language.
If tf was made by someone else, I would understand the choice, because the popularity of python and tons of library available, but since Google has unlimited resource, and AI is clearly the future, I really expect they have the courage to choose something else.
Sometimes I wonder, AI may gone rouge some day - not because that we deliberately make it so, but that a bug somewhere in our code.
2. Lack of maintenance.
We all like shiny new ideas, and get excited implementing them, but once the fun part is done, so goes the excitement. A good library needs to be tweaked and re-tweaked, some of these need boring hard work, smart people don't like that.
But hey, Google is doing this for free, as long as it's not deliberately made so (to stall the community), we should be appreciate, it's a open source project and that don't just mean we can use it for free, but also that we should done our part to make it better.
Lots of brilliant people working on heavily resourced projects ... but also significant bureaucracy and many political animals in what was formerly a pristine engineering "garden of eden".
You can see a lot of Google projects struggling now, and many startups in the same space as Google projects doing much better than more resourced teams doing the same thing at Google.
Nobody ever got fired at Google for spending all day brilliantly arguing on Google's internal newsgroups and not doing any real work. And it shows in Google's work culture. Imagine a person doing that at a startup ... or Amazon for that matter. They would not survive very long.
If you are young, and have many years of productive/earning years ahead of you, it might make sense to turn down a Google offer to try something a bit more "bloody" and hectic for a few years before you settle down in a comfy Google job.
There are many brilliant people you can learn from at Google ... but very few work very hard. And hard, productive work is a skill to learn too.
There was a pretty high profile example of this not all that long ago.
In fact, to do even half decent export of TF models, you have to switch to keras to try and do any kind of export.
I have a 10 email conversation with enterprise Google Cloud support to try and get a ML Engine output serialised to work on Android.
There are threads open all over the place on stackoverflow and elsewhere - and yes, we have tried all SIX ways.
Admittedly, the documentation in this area is extremely bad and I basically had to figure out myself how to do it, though this was long before 1.0.
That said, would you be able to share any example snippets on how you are persisting and loading these models in your code ? That would be super helpful.
Also, im getting the feeling that you are using the deprecated method of saving. I think they are shifting to Metagraph now (not sure about this) https://www.tensorflow.org/versions/master/api_docs/python/t...
[1] https://github.com/tensorflow/tensorflow/issues/10254 [2] https://github.com/tensorflow/tensorflow/blob/master/tensorf... [3] https://github.com/tensorflow/tensorflow/blob/master/tensorf... [4] https://github.com/tensorflow/tensorflow/issues/10299
The relevant code is (still) in private repositories, but I have a deck with some examples using Rust:
https://www.dropbox.com/s/t8r056f6wqlktqv/embedding-tensorfl...
The only effective difference IMHO is that the declarative style does deferred evaluation which makes it more difficult to examine "intermediate states" (aka "debugging") but as the author stated themselves, this is simply a matter of outputting those intermediate states/setting them as outputs... this is no different from examining intermediate state in every programming paradigm ever invented. Unit-testing every step is IMHO a far better way of debugging/ensuring validity, but this is possibly much more difficult in the "black box" environment of self-weighting neural networks than it is in your more traditional programming environment
> Now let’s look at a Pytorch example that does the same thing
Reminds me Escobar (russian philosopher, not to be confused with Lopez-Escobar) theorem, which approximately translates into English as "With no alternative choice of the two opposite entities, both will be an exceptional nonsense."
I really hope at some point this entire universe gets liberated from Python at some point. Even R would be more palatable. Both these examples are awful, error prone, obfuscated, and beholden to Python's difficulties with large sums of data.
I tried to do so myself and couldn't come up with anything significantly better, but I've been writing Python for a long time and might just be stuck in a local minimum :)
In julia
Another comment has already pointed out that Julia framework. I'll relink it for completeness: https://fluxml.github.io
You can also look at Haskell's Grenade examples: https://github.com/HuwCampbell/grenade
Fundamentally different because the notion of what the compiler should be doing is fundamentally different, and the notion of how data should be input is somewhat different.
But I'd even take Scala at this point, and you could do worse than Clojure.
Python code isn't especially "ugly" as in line noise. It's ugly in the sense that it's a deluge of outdated and ineffective programming paradigms that maximize the difficulty of writing code.
It's early days but so far F# it ticks all my boxes and I'm finding it an extremely nice language to write in. It is just as terse as Python (due to type inference) but statically typed, and there is a great plugin Ionide for VSCode which makes for a really polished development environment.
Plan is to use Microsoft's CNTK for ML/DL stuff.
For anyone frustrated with Python's duck typing, I highly recommend you check out F#.
But it is productive to say, "An excessively declarative style has been considered bad in programming for over thirty years."
That is reminding people this is a largely solved problem that the industry refuses to embrace because our legacy rube-goldberg contraption of software refuses to accept ever abandoning anything in favor of simplicity.
A great example: Why are both of these examples using the error prone pattern of explicit bounds iteration?
Ummm. No. 'Objectively' is utter nonsense. For an objective view we would need to define "better" first and measure both interfaces performance. I think it is preference. I prefer the Tensorflow interface and don't mind it's declarative style.
However, if one wants to criticize something one could start with the static nature of Tensorflow (which you rightfully mentioned), which makes it hard to do stuff like LSTMs and dynamic batching. That works better in Torch. To me that is the only real attack point for Tensorflow. But keep in mind, that Tensorflow is still "1.x" software and stuff like Tf-Fold addresses the problem with static graphs already and Google plans for more dynamic graphs in Tf 2.0.
Also it would have been nice of you to measured performance of the both frameworks on common problems. But looking at the interface and shouting "bad" at Tensorflow is not really critique but a personal dissatisfaction with Tensorflow.
I think that Tf currently aims more at production code than research stuff. The mentioned problem with being to low-level for simple stuff like layers is also not right. Have a look at the shipped contrib modules, you'll find common layers in there.
That said, you’re right. There’s no way I’d deploy it to production.
Additionally, PyTorch download page warns you point blank that it’s an early version of the software and that you should “expect some adventures”. Adventures are fine for research, but inadvisable in production IMO.
Making it run at scale using something like kubernetes is more advanced stuff but still within good devops practices.
Unless, of course if you want to run on exotic HW, like TPUs. But that's an issue for Googlish scales.
Edit: BTW, if on premise you mean Windows client software your point is totally valid, Python would suck for this.
Re: Windows, there's no TF serving for that, so TF is no better than other Python frameworks.
Disclosure: I work at Google, but not TensorFlow.
You could mean large scale, real time, "small batch job with online lookups from a CRUD database",..
Let's just admit it's use case specific and move on.
I am not saying that Pytorch is bad in production but I fail to see how your metric of 8 images per second proves anything.
Disclaimer: I work on Caffe2 team (not on ONNX, though)
I'm super skeptical of these interchanges, because it seems very difficult to avoid train/test skew. Any difference in detail between the two implementations is a potential problem. I can imagine different order of operations putting out some values by 0.1%, causing 1% eventual loss of accuracy.
Implementation details of course differ between frameworks, but luckily neural networks are very robust to noise. In my experience, changes like using Winograd/Fourier for convolutions, or even running the whole thing in FP16 do not result in noticeable artifacts, and these are among the biggest differences you could have between frameworks.
I've seen this stated by a few people, but never justified. What are some reasons for not deploying PyTorch in production?
IMO, learning TF or pytorch is more effective at least in the current state of affairs.
* If you are implementing a standard model (that's 90% of industry use cases, and a large fraction of research use cases as well), Keras primitives considerably simplify your workflow and make you a lot more productive.
* When you need to implement something highly customized or unusual, you can revert back to writing pure TensorFlow code, which will integrate seamlessly with your Keras workflow (via custom layers, functions etc).
Basically, Keras increases your productivity for common use cases, without any flexibility cost for rare/custom use cases. It is meant to be used together with TF, not as a replacement for TF.
import tensorflow as tf
import numpy as np
X = tf.placeholder("float")
Y = tf.placeholder("float")
pred = tf.layers.dense(X,use_bias=false)
cost = tf.losses.mean_squared_error(labels=Y,
predictions=pred)
optimizer =
tf.train.GradientDescentOptimizer(0.01).minimize(cost)
with tf.Session() as sess:
sess.run(tf.global_variables_initializer())
for t in range(10000):
x = np.array(np.random.random()).reshape((1, 1, 1, 1))
y = x * 3
(_, c) = sess.run([optimizer, cost], feed_dict={X: x, Y: y})
print cthe android SDK appears to lack a coherent, overarching concept or set of guiding principles which the client programmer can internalize and rely upon. there are many special cases and inscrutable behaviors.
like, sometimes you ask for a piece of work to be done by the SDK and it calls you back on your implementation of a Listener class. but, other times, you have to register a BroadcastReceiver. still other times, you have to override OnActivityResult. but, then, there are these other times when you have to create and provide a PendingIntent.
you go through enough of this stuff and you start to wonder: "Did there really need to be soooo much variety in the way these SDK methods hand info back to the client? Couldn't they have standardized this?"
it's almost like Google turns their (very smart) programmers loose on the SDK and never looks back. google seems to just defer entirely to their opinions and judgments. smart as these programmers are, they seem to have different tastes, different approaches to API design.
and it all goes into the SDK.
Scikit-learn however doesn't do deep learning. That being said most problems faced by mainstream tech workers don't actually have huge data sets or need deep learning and for those problems scikit-learn is great. Tensorflow only really comes into its own if you have/need some combination very huge data sets, very deep networks, a novel or non-standard network configuration and a large cluster of machines to run your learning on. If you want to apply a standard ML algorithm in the standard way on a 'small' dataset and aren't super constrained by performance then scikit-learn is almost always a good choice.
Why can't we just have compute networks that can be used for anything, from computational linear algebra to deep learning?
> Let’s be honest, when you have about half a dozen open source high-level libraries out there built on top of your already high-level library to make your library usable, you know something has gone terribly wrong
I consider that not a bug but a feature.
You mean, like Tensorflow?
CuPy is even designed to be drop-in compatible with Numpy; I don't think there is much support for LAPACK routines at this point though.
Doing MATLAB-esque work with tensorflow would require some special tolerance for pain, or would require one to be a masochist.
http://dynet.readthedocs.io/en/latest/index.html
If you're doing the sort of work that would be published at ACL or EMNLP, DyNet is a really good choice.
Rad! Do you have any examples (or literature) that explains when this is beneficial?
[1] http://papers.nips.cc/paper/893-a-growing-neural-gas-network...
- Bloated build system that is near impossible to get working - who even uses maven ?! Pytorch/Caffe are super-simple to build in comparison; with Chainer, it's even simple: all you need is pip install (even on exotic ARM devices).
- The benefits of all that static analysis simply aren't there. In addition, PyTorch has a jit-compiler which one can argue lets one have their cake and eat it too.
- Loops are extremely limited. Okay, we know RNN/LSTMs aren't really TF's thing, but if you venture out to do something out of the ordinary even making it batch-size invariant is difficult. There isn't even a map-reduce op that works without knowing the dimension at compile time. You can hack something together by fooling one of those low level while_loop ops, but that just tells you how silly the whole thing is.
> Pytorch’s interface is objectively much better than Tensorflow’s
How does your subjective opinion about Tensorflow suddenly become objective? I don't see any fact based or metrics based argument that establishes this objectively. All I see is an argument based on your needs and preference. That is far from objective.