Show HN: End-to-end deep learning experimentation platform for Tensorflow
github.com
github.com
The first homework assignment was image recognition on notMNIST. That dataset has been used thousands of times before, so I decided to make it more interesting! I got data for 1500 Chinese fonts, so I could make a better OCR for Chinese characters.
I put that into the example, made the "pickle" files, trained the model, got a 90% prediction-test result value... but now what?
How do I take that model and make a real image OCR that can take a photo from my phone and analyse it?
I can write scripts to get data, and follow the steps in a course to make a model, but unless I have a useful output, I don't know how I'm going to apply this to my other programs. If you're offering "end to end", then documentation and examples of some kind of API would be great. (e.g. send a PNG to this REST service, and get the prediction back as text).
I assume you're talking about some kind of endpoint to submit the image to, and offload the computation entirely.
However I just wanted to clarify for people that are new to this, that running the neural net on a mobile device usually requires entirely different models, because of the memory/compute limitations. It's a rapidly developing field, but it is not quite as accessible as developing for desktops yet.
I do believe Tensorflow allows you to run models on mobile now though [0]
Real life doesn't give me folders of training data to sort. It gives me photos and videos.
Applying a systems engineering approach, I think I need to:
1. Break the picture down into thousands of small little squares.
2. Look for a character inside each picture.
3. If a character is found, add the character to the output text.
That brings new questions. How big should each square tile be? What if the characters are not perfectly flat? How do I avoid getting duplicates from tiles that are next to each other?
Whether the tile->text conversion is done with a simple classifier or "deep learning model", the bulk of the work is done using non-Tensorflow programming. I don't know about open-source projects I could build on and retrain using my data. So the most likely result is that I'll show off the fancy Tensorflow data as a "portfolio project" and archive it, never to be used in production.
That's interesting. Could you expand on this some more?
thank you for taking the time to have a look at the project, and I am very happy that we are receiving such a constructive criticism.
The way to create a TF record is still manuall, here's an image data converter that was used to create the mnist dataset:
https://github.com/polyaxon/polyaxon/blob/master/polyaxon/da...
More data converters will be available soon.
Once you have a record, numpy array, or a pandas Dataframe, you can basically use any operation/layer on your data by providing the feature name and the list of operation to apply. In general only operation that are necessary for data augmentation should be done on the input data pipelines, otherwise everything should be done during the creation of the TF records, to minimize the computations.
One last thing for reinforcement learnign, the way we feed data is through the feed_dict, because an interaction with an environement is necessary.
For the same reasons, I have nothing against people who "do anything that resembles anything that has been done before," and I certainly wouldn't attempt to discourage the author. I'm sorry if that cartoon is over-referenced (in attempts to be negative), I can't really control that, but I think it's something to think about when starting a new project (i.e. what is your goal, does it differ from other projects, is that important to you?)
Sorry if my comment seemed like I was shooting this down, I certainly had no intent of doing that.