TensorFlow Benchmarks
github.com
github.com
DNNs provide some interesting opportunities for domain specific optimization with lossy math because empirically it seems that the solutions dont require significant precision. Something like a general Tensor library holds you back from using some of the tricks like synchronization-free hogwild, where it wouldn't be appropriate in other scenarios where more exact computation is necessary.
One interesting step towards this is their implementation of lossy compression with a custom 16-bit half float representation. However a lot of people are dialing down the precision even further, including one of the authors of TensorFlow (http://petewarden.com/2015/05/23/why-are-eight-bits-enough-f...).
Scott Gray, one of the principle engineers at Nervana Systems, has a very wise reply in the github thread - they aren't doing much to optimize for memory locality, and wont have amazing speedups until that happens.
They must have better optimization a for running in production, such as in place operations.
For its predecessor DistBelief, here are some quotes from NIPS 2012 paper:
"The moderately sized speech model runs fastest on 8 machines, computing 2.2x faster than using a single machine. (Models were configured to use no more than 20 cores per machine.) Partitioning the model on more than 8 machines actually slows training, as network overhead starts to dominate in the fully-connected network structure and there is less work for each machine to perform with more partitions." ("The moderately sized" model here has 42 million model parameters. Check the paper for details.)
"In contrast, the much larger, locally-connected image models can benefit from using many more machines per model replica. The largest model, with 1.7 billion parameters benefits the most, giving a speedup of more than 12x using 81 machines. For these large models using more machines continues to increase speed, but with diminishing returns."
Can anyone here present a use case where they have personally needed horizontal scaling because a Titan X couldn't fit what they were trying to do? It's the thing I hear talked about the most and used the least.
The biggest misunderstanding I've heard is "I have petabytes of data so I need multi-GPU", but NNs are trained on batches of data, so it doesn't matter. It's the model that must fit, but 12gb can fit a big model.
The whole "who can build the biggest net" is like building the fastest car: it's for show, the slower ones are way more cost effective and get you there just as well.
Keras is my present deep learning library recommendation--it's like torch 7 but in python and theano. Multi-GPU is roadmap because there's a lot more going on in deep learning that acts as a bigger differentiator than horizontal scaling.
For the record, when I say scaling, I'm not talking about having multiple machines serve multiple requests (that's a good reason to purchase more machines), I'm talking about spreading your gradient computations and/or model weights across machines or GPUs.
Data parallel is easy to implement and leads to linear training time speed ups, as per 1 week to 2 days with 4x hardware. But not bigger models.
Model parallelism leads to bigger models and is what I've been referring to. It is overkill and doesn't work well anyway. Frameworks should not be expected to support it because there are more interesting topics in deep learning they could support instead.
Better research/methods will come out at some point, at which point this calculus will change, but not yet! Today it is the very definition of premature optimization in nearly all cases.
Of course, there are no clear wins (generally, losses) in computation by scaling RNNs horizontally on GPUs - but sometimes you really do need more than 12GB.
No clear wins in horizontal scaling
Reduce batch size
Use truncated back prop
*search for better hyper params*
I usually do 2-4. After those, have you really seen scaling result in significant accuracy gains? And what percent of the time is that necessary? Genuinely interested--and you guys rock btw!It's difficult to say what "most applications are" but many algos currently seem bottlenecked by how fast you can load data into your GFX card (hint: it's slow). So what you're saying about batching is correct, but it does matter.
All of them are, but that's a PCI-e issue and horizontal scaling doesn't fix that (unless using nvlink or similar afaik, but then you face the fact that current horizontal scaling schemes aren't very effective at increasing model accuracy anyway)
> It's difficult to say what "most applications are"
Nevermind "most applications"--so far, all I've heard is one, that being the absolute bleeding edge of RNN research, assuming you're using a huge softmax instead of an alternative.
My point remains: Multi-GPU is way down the list when it comes to features a DL framework should have. Because very very few people need it.
kajecounterhack, do you use multi-GPU for your DL work? If so, how often?
I'm using it right now, to great effect. I can't really say what for, but methinks I'll be using it every hour of every day for the foreseeable future.
But IMO the real obstacle to horizontal scaling is the communication between servers not the usefulness of doing so. Within a server, one can pack 8 TitanX/M40 GPUs with high speed 13 GB/s P2P communication between them (and up to 16 GPUs in various unproven science project servers). That's ~50 TFLOPs and 96 GB in a box. That rocks. Just ask the guys at mindori (if they ever ship that is).
But between servers lies a sippy straw of 100 Gb/s Infiniband at best or worst-case, ~1 Gb/s on AWS GPU servers with freaky nearest neighbor weather. If you can't make efficient use of 8 GPUs in a single box, I agree, don't bother breaking out to the next server.
That said, frameworks like mxnet have opened the floodgates to experimenting with larger models and more distributed training algorithms. Time will tell if this pans out. But 2 years ago, Andrew Ng's group showed 12 GTX 680s distributed across 3 servers kicking the crap out of Google Brain. I expect more of this, not less, in the near future.
The embeddings for text models with a real world vocabulary can get very very large, especially if multiple languages are involved.
But assuming you're serious and not dabbling, you can build a ~27 TFLOP server for <$7K in parts or buy it assembled for ~$9K.
http://exxactcorp.com/index.php/solution/solu_detail/225
Or if you just hate the thought of money burning a hole in your pocket, pay NVIDIA >$15K for the same thing with a pretty NVIDIA logo:
https://developer.nvidia.com/devbox
If you're dabbling, just go buy a $1000 TitanX and you're instantly a minor deep learning superpower.
I think the key to TensorFlow is not how fast it runs on 1 machine; but how fast it runs on 10,000.
Consider map-reduce (Hadoop). Sure, you can sort 1GB data on a single machine 10x faster (using /usr/bin/sort) than using Hadoop on that machine; but make the data 1TB and add 1000 machines, now lets see how fast you can sort with /usr/bin/sort!
It's also getting out of hand because everyone has already decided TensorFlow must be amazing, so everyone is extolling the virtues of what right now we only know to be speculation and PR.
A position of authority doesn't mean the person is immune to jealousy. Someone in a high status position is going to be much more likely to have a knee-jerk defensive reaction to news that threatens their image.
These benchmarks evaluate single-node performance. LeCun's remarks were concerning distributed training (specifically that bandwidth between machines is a limiting factor to scalability) -- which we can't test yet since the current version of TF is single-node only.
Dean's response in the video "it depends on your [computer] network" is an interesting response :).
If TF really has a way to make general, networked, distributed training efficient (more than 1 bit weight updates, low precision weights and all the other crazy tricks which already exist) - that is truly remarkable and they rightly deserve huge kudos.
If they are faster in distributed training only on Google machines or with Google's network architecture, that isn't really a useful datapoint for the general public.
We could wire everything with 10GB/s NiCs, change to jumbo frames, pull all these other bandwidth reducing tricks, trick out the Linux kernel, etc. and then your network probably won't matter - but that isn't really general or cheap. The key will be what is the minimum effort necessary to avoid network bottlenecking, and does TF improve that minimum level over existing solutions?
It's more realistic to discuss librairies for 1 machine / GPU
Still faster than Hadoop, because sorting the data will be faster than sending it over a network. What should I do with the other 999 machines?
(I know what you actually mean, by the way. Make it 100 TB and say that the data is already spread out across 1000 machines.)
It is not surprising that a tool developed and focused on `Google scale` work has some imperfections in a wildly different setting. The question is - will they (or some dedicated contributor) speed up a use case the business itself may not have a use for? My gut feeling is that they will, but these things usually don`t happen overnight. Torch, Theano, and Caffe have years of work put into them and have largely been focused on the one to two GPU case.
NVIDIA support alone will make sure TF knocks it out of the park down the road. IMO it's crazy to consider the first release the final say on TF's GPU performance. Caffe, Torch, and Theano have a huge head start (and lots of pre-existing technical support from NVIDIA).
The biggest limitation I see right now is that their multi-GPU algorithms are really simple and inefficient. That will change I'm sure now that they're getting benchmarked against everyone else.
Even with v4 support which is coming soon, the general setup that users have is one to two GPUs in one machine. This means adding the right inplace operations can have a huge impact on performance, and is probably a usecase that Google has not focused on given their internal infrastructure. I am sure they will normalize to "at or slightly above" the performance of other toolkits, but the question is when?
This benchmark doesn't have anything to do with multi-GPU as far as I am aware - these are single machine, single GPU results. I would wager 90% of the deep learning hobbyist and research communities run in this setting, so benchmarking this is really important.
For people with huge amounts of networked and distributed resources - distributed support will be very amazing!
No surprises there whatsoever. The tensorflow engine is the most gold-plated POS I've seen in a long time. if it were running on 1000+ servers, I'd get the level of overengineering they've applied here. But single server pthreads? WTF?
Also parameter server is dumb unless you sweat the implementation of the gather/reduction ops, but I digress.
I've been running convnet-benchmarks forever now, and I've been running them independently on separate personal hardware. I do this as a hobby.
I've done an apples-to-apples comparison, and my benchmark review only puts the facts forward, I dont attack them.
If you read my other social media comments, I've been pretty positive about TensorFlow, and I've even put in some groundwork to write Torch bindings for it.
Please stop spewing nonsense interpolated from like 2 super-weak data points.
The responses seem to show that the way you implement things can make a big difference in runtime. Perhaps the scripts used for benchmarking can be further optimized?
That said, the lack of in-place operations might be surprising (although it has been said that they are coming)
a) It's Google, they can throw hardware at a problem such as out of memory
b) TensorFlow makes methods development so much easier that it's worth the loss of performance
c) It's early days and the compute graph scheduler has lots of opportunities, and is designed, for optimization, and in a more flexible fashion than other frameworks.
When I work on developing methods for scientific code, I worry more about whether the code is bug free/easy to understand and that it's giving the right answer. I don't usually worry about performance unless I'm actually not able to run things. Especially since when developing stuff you waste way more time on runs with bugs - I dread to think what my (published / unpublished) CPU hour ratio is. If the new approach allows less buggy implementations then that's a resource win.
Indeed, if TensorFlow means I can try out an idea with 1 day of coding and 2 days of training rather than 3 days of coding and 1 day of training then I can spend a day drinking cocktails and reading books and still be finished sooner.
Based on the tutorials it seems like I'd be able to pretty quickly build a translation pipeline, and in fact there's an implementation of that I think I'll try. If it takes a week or two to train, that's fine by me, I've got other things to be getting on with.
"Hey, we would throw this away, but maybe this could be used to shine up our brand a bit"