Simple example of machine learning in TensorFlow
github.com
github.com
For your next tutorial, may I suggest: 1) a list of do's and don'ts for constructing a savable/restorable model, and 2) a wee bit of example code.
Of course, now that I have discovered Keras I'm moving away from low-level direct TensorFlow. But I suspect I'm not the only one a bit foggy about the whole save/restore work flow.
I find that saving and restoring are of the weirder things with TensorFlow, you can either go all out an decide to save out all the variables, or only the ones needed for the model.
You usually don't want to save out gradients (which are also variables) since they take up a bunch of space and aren't actually that useful to restore. Now on the other, what are model variables -- do you want to save model variables + the moving averages ... or just the averages. But then when you're loading you'll have to "shadow" the moving averages to the real variables that actually run in your model.
Good news though, most of the scaffolding code you can write once and re-use it over and over again.
Models with a few million parameters result in a file around ~50MB, which is still reasonable for modern production use cases.
https://blog.metaflow.fr/tensorflow-saving-restoring-and-mix...
I think that having to use a duck typed language makes it so much harder than it has to be. The API is huge, and you get really no help from the IDE. If I could do tf.train.GradientDescentOptimizer as in the example, and then get autocomplete to what I can do next with that object it would really help. Or list functions that takes that kind of object as input. Now, one is searching in the dark.
However, Estimators are not in TensorFlow core, which means the API isn't fixed quite yet.
This blog series was also helpful on a conceptual level: https://medium.com/emergent-future/simple-reinforcement-lear...
LOL :) (Side-note: 8 million is still not big data)
lol
lol
I get the output of the model (y_model = m*xs[i]+b), it's the y = mx + b where we know x (from the dataset) and have y be a variable.
The error is where I start to lose it, so I get the idea of the first part (ys[i]-y_model). It's basically the difference between the actual y value (from the dataset). I get that we want this number to be as small as possible as the closer to zero it is for the entire dataset that means we get closer to the line going through (or near) all the points and the closest fit will be when this total_error is nearest to zero.
What I don't get is the squaring of the difference. Is it just to make the difference a larger number so that it's a little more normalized? How do you get to the conclusion that it needs to be normalized? Same thing with the learning rate? I believe these to be correlated but I can't tell you how...
sum_errors_A = 4 + -3 + -1
sum_errors_B = 1 + 1 + 1
B is obviously the better model, but it has a higher error than A when comparing. If we squared all the terms and then added, B would be the stronger model.
The following answer is about standard deviation but the desired properties from squaring is also relevant to your question:
http://stats.stackexchange.com/questions/118/why-square-the-...
There are also historical reasons for using squared error. The square function is smooth and differentiable to you can analytically solve for the gradient. Before fast computers this was crucial for solving regression problems as a closed for makes everything easier.
I currently have a small pet project where I think some simple ML would be cool but I don't know where to start so these things are great.
Basically my use case is that I have a bunch of 64x64 images (16 colors) which I manually label as "good", "neutral" or "bad". I want to input this dataset and train the network to categorize new 64x64 images of the same type.
The closest I've found is this: https://gist.github.com/sono-bfio/89a91da65a12175fb1169240cd...
But it's still too hard to understand exactly how I can create my own dataset and how to set it up efficiently (the example is using 32x32 but I also want to factor in that it's only 16 colors; will that give it some performance advantages?).
Here's how I solved a similar problem in my deep learning class. Instead of classifying between 10 animals you would just have 3 possible labels.
https://github.com/hermiti/deep_learning_project_2/blob/mast...
Perhaps you were meaning to put "bare bones"? Google's definition of the latter is "reduced to or comprising only the basic or essential elements of something."
Don't want to detract from your point but I think your title is throwing some people off. I know I would be hesitant to click something at work that sounds like it could contain nudity.
"The slope and y-intercept of the line are determined using gradient descent."
What on earth does that mean? Maybe they should teach mathematics in english at universities outside of english speaking countries. German mathematics does not help here.
I wish there was a 4GL like SQL for machine learning using dynamic programming for algorithm selection and model synthesis like a dbms query planner.
PREDICT s as revenue LEARN FROM company.sales as s GROUP BY MONTH ORDER BY company.region
Slope and intercept are very standard names for the parameters of a linear regression model. Gradient descent is the name of the algorithm used.
and it is absolutely a good idea, as long as you include validation and QA abilities right along side train/predict.