TensorFlow – Consise Examples for Beginners
github.com
github.com
It's always left as a useless exercise for the reader to divine how to generate such a dataset from his/her own data.
Examples should be way more general. The starting point shouldn't be:
from tensorflow.examples.tutorials.mnist import input_data
mnist = input_data.read_data_sets("/tmp/data/", one_hot=True)
, it should start with "here is a directory of images and their classes" and end with a CNN model.EDIT: Should anyone have any insight as to where I might get such a tutorial (or have the desire to write one), I know a herd of ML pre-initiates that would be grateful.
Read the how to on reading data from files: https://www.tensorflow.org/versions/r0.8/how_tos/reading_dat...
and check out this useful Stack Overflow answer: http://stackoverflow.com/questions/33648322/tensorflow-image...
The other problem here though is corpus layout. Image net and the academic datasets typically require special readers, but even then: for actual datasets there's a few ways to do the corpus layout. In our experience there are 2 things that have worked well have been: a folder per label or labels in the name.
Then there's still having balanced minibatches though. When you get to having balanced minibatches images segmented by folder means you don't get balanced minibatches out.
Then there's disk to think about, do I really want to re run the same pre processing every time I train? Then: how do I explain that to a beginner?
So I'm probably going to want to have a corpus generator where we end up with a pre saved/balanced minibatches for training..which leads us back to what you see now.
A good middle ground here might be a corpus generator that takes all the minor stuff like that in to consideration...but still data is messy.
I would suggest looking at the wealth of imaging libraries out there in python and building something based on a "from scratch" image corpus.
Minor plug: We thought about that a lot in building deeplearning4j. http://deeplearning4j.org/canova
You may not use java but the idea of "vectorization" is still a good one I think any ml practitioner who's touched pandas could appreciate. We built an abstraction called a datasetiterator which auto magically returns the batches for people so they don't have to think about the details but still having access to "real" data. I'm not sure what the python equivalent to this would be though.
A good exercise would be to figure out how to extract the data and put it into a numpy array. Then you can test on most - if not all - of the frameworks.
Say you need a CNN text classifier algorithm to categorize simple single page documents. So you set one up via TensorFlow, train it with a big dataset, and get it outputting categories with decent accuracy. Could you then use some type of API and query it from a (low traffic) production web app?
Or is it more for the research phase rather than real-time interaction?
[1] https://medium.com/@sentimentron/faceoff-theano-vs-tensorflo...
[2] https://github.com/tflearn/tflearn/blob/master/examples/nlp/...
https://github.com/soumith/convnet-benchmarks
They place TensorFlow performance on par with Torch (within 10%).
TensorFlow is like Numpy, only it is capable of working symbolically and can work more easily with your GPU in a highly parallel manner. For those two reasons, it is vastly superior to numpy for tasks like deep learning / machine learning.
Theano is another symbolic numerical library, coming before TensorFlow, though TensorFlow has seemingly gained more popularity and people have more general faith in it.