Neural networks in JavaScript
github.com
github.com
Just out of curiosity, does anyone have a link on NN training parallelization? I’m just hitting paywalls. I’m curious if it’s simple or if there are dependencies between neurons that make it hard. As I see it, the next big hurdles after ubiquitization are distributed training (possibly with GPUs) and empirical data on the best tuning parameters for various problem spaces (which will come largely for free when training happens 5 or 10 orders of magnitude faster than today). Eventually we’ll have an answer as to whether AGI can be composed of basic building blocks or if we’re still missing some fundamental insight.
For example, genetic algorithms are fairly easy to parallelize because by their nature they use large and independent populations. About the time I was graduating college in the late 90s, there was great interest in using them to find the initial weights in NNs to avoid local minima, but then I didn’t hear much about it since. Have there been any breakthroughs?
Geoff Hinton popularized pre-training nets with unsupervised latent-feature learning autoencoders in the 90s, making training nets much faster. More recently GPUs have been used because they perform matrix/vector multiplication much faster. Computer clusters of hundreds and thousands of nodes are used to run SGD at high speeds with the above techniques, and pretraining with autoencoders isn't quite necessary anymore.
Bayesian optimization is one technique which addresses your concern regarding parameter-choosing. Essentially Bayesian optimization uses Bayesian reasoning to plan experiments which seek to discover the most information about the loss function of the network, thus allowing you to efficiently choose hyperparamters of your model (whatever it is).
Neural nets have definitely been seeing breakthroughs since the 90s–the last few years alone have been amazing. Check out "Google Brain", "ImageNet neural net", "Never-ending learning (NELL)", Geoff Hinton (reinvention of nns), Tom Mitchell (online learning, proto-AI), Yan LeCun (convnets), Nando de Freitas (bayes opt). Lots of good stuff out there!
Also, I can't use backpropagation for what I'm trying to do.
Thanks for the suggestions. I had heard of Yann LeCun, I'm trying to do image recognition in an "different" way than CNNs, I'm going to read up on the other people in the list.
I can't use backpropagation because some of the output nodes are used to apply a transformation on the input followed by reapplication of the (now transformed) input through the neural net.
The approach called iterative reduce is basically training on mini batches in parallel and averaging the outcomes. The trick is to do it in multiple iterations on each mini batch averaging the parameters at runtime. It's a surprisingly simple technique that works really well.
Unfortunately, it's not as async as sandblaster LBFGS by dean and co[3].
Andrew himself has told me that the gradient was unstable with this approach though.
Right now, distributed training of neural nets is still like the wild west. I can see a few approaches coming up over time though.
[1] https://www.youtube.com/watch?v=h2Ixpfn-DTg
[2] http://deeplearning4j.org/
[3]: http://research.google.com/archive/large_deep_networks_nips2...
> the same way we learned the C++ STL back in the day
My school forbid you from using STL. You had to build all of your data structures in freshman/sophomore years and then use those same data structures you built in your sophomore/junior/senior projects. It was really crazy because if you never got your lists, trees, graphs, and hashes working you basically wouldn't make it through the CS program. It sucked at the time, but I'm grateful now.I've tutored students in Java classes who would write a bubble sort every time they needed to sort a List, instead of using Collections.sort, simply because they didn't know Collections.sort existed! Or their own ugly code of looping, indexOf and substring to get values from a CSV instead of using string.split(",") or a premade csv parser (which I know that professor would have been fine with, he would have even encouraged it).
Unless point of the project is to learn how to make one from scratch, you shouldn't be wasting your time writing an alphabetical sorting function when you can be solving novel problems.
> Wow, that sounds very bad to me.
It might, but today, I run into tons of developers that don't understand the fundamentals. I worked with a guy who I affectionally named "Stack Man." He basically used a stack for any list implementation. He didn't understand the difference between a list, a queue, and a stack.You're teaching your students Java and telling them "Just use Collections.sort" and that's fine, but do you understand when Collections.sort is bad? If you don't understand the underlying algorithm of how a merge sort works, you might be telling them to use something that is worse than could actually be worse than a bubble sort in certain cases.
Also Java wasn't taught to CS students at our school, that was only for BIT (Business Information Technology). Teachers wanted you to understand memory and pointers.
That's brilliant, although I'd imagine it could begin to annoy. Slightly obtuse, but I'd have had to get a snippet of this to play every time he checked in a new stack...
When your assignment has to be uploaded in 3 hours, and your dataset is 120 items, Collections.Sort is always better than writing your own sort.
I never said I taught them only to use Collections.sort, as that would be just as bad as never teaching them the standard libraries and only teaching data-structures. You need a healthy balance of both fields of knowledge. Otherwise you end up spitting out either programmers who can use library functions without understanding what's actually going on, or programmers who know all about the underlying methods and logic, but can't write code efficiently and quickly enough to actually help a team.
That's not to say you cannot write code efficiently or quickly. You likely did as most good programmers do and learned on your own time as well. But for students who don't do this, they will be lacking huge amounts of practical knowledge.
If you're trained as a salaried code churner to solve some "pragmatic enterprise problem" yes.
If you're trained as a computer scientist, you have to know how to do stuff properly from first elements, and also how to invent and code similar new stuff, custom tailored to new domains yourself with ease.
You don't get progress in computer science with people just learning to use existing libraries and data structures. At best you can get some innovative apps, but not people able to build the substructure to create the innovative apps of tomorrow.
Not to mention the sorry state of most existing data structures libraries and core APIs, from STL to the Java SDK. A lot of this stuff needs to be burned with fire to be brought to 2014 levels.
Off the topic of neural networks in javascript but on the topic of neural networks in general, the readme links a cool paper on applying convolutional neural networks to old games.
I added an extra input to the training data for the XOR example:
net.train([{input: [0, 0], output: [0]},
{input: [0, 1], output: [1]},
{input: [1, 0], output: [1]},
{input: [1, 1], output: [0]},
{input: [3, 4], output: [7]}])
But the result of running: var output = net.run([3, 4]); // [0.987]
Is [0.9999823694382457]
which strikes me as odd. Shouldn't it be closer to 7? Each training pattern should have an input and an output, both of which can be either an array of numbers from 0 to 1 or a hash of numbers from 0 to 1.Also keep in mind that there needs to be neurons to convert a value into binary, and back! Which is also going to take a lot of neurons, you could try doing dec -> binary conversion before sending it to the neural network:
net.train([{input: [0, 0, 0, 0, 0, 0], output: [0, 0, 0]},
{input: [0, 0, 0, 0, 0, 1], output: [0, 0, 1]},
{input: [0, 0, 1, 0, 0, 0], output: [0, 0, 1]},
{input: [0, 0, 1, 0, 0, 1], output: [0, 0, 0]},
{input: [0, 1, 1, 1, 0, 0], output: [1, 1, 1]}])
var output = net.run([0, 1, 1, 1, 0, 0]); // [ 0.919125109533231, 0.9195887654595207, 0.9734227586511985 ]https://en.wikipedia.org/wiki/Logistic_function
If you want to create a neural net that can, say, identify handwritten digits and output a result 0..9 it is usually better to create a neural net that has 10 outputs, each corresponding to one digit, rather than a neural net that has one output, say, parsed between (0,0.1) = 0, [0.1,0.2) = 1, etc.
In my case this was a big ol' row of titties, right in the middle of my office..
http://en.wikipedia.org/wiki/Feedforward_neural_network#Mult...
You can interpret boolean functions as classification problems: given some input(s) x, classify the output as either 1 (class A) or 0 (class B).
The truth table of an XOR function looks like this:
x1 | x2 | XOR | Class
---------------------
0 | 0 | 0 | B
0 | 1 | 1 | A
1 | 0 | 1 | A
1 | 1 | 0 | B
If you were to plot the XOR function in R2, it would look something like this: x2 ^
|
|
A B
|
|
B-----A----->
x1
To be non-linearly separable means that there is no straight line that you can draw in the above plot that will split the classes. Thus, we need a non-linear classifier, e.g. a neural net.As in, the following would be linear separable (all labels on the same side of the line)?
x2 ^
|
|
A A
|
|
B-----B----->
x1In this simple 2-d case the subspace is a line, so this example is linearly separable. In the 3rd dimensional case linearly separable refers to classification using a single 2-dimensional plane. For 4 dimensions with the data, you need a single 3-dimensional hyperplane for it to be linearly separable, and so on.
v(x1,x2) = C1 * x1 + C2 * x2 + C3.
C's being "coefficients".
In your case, v(x1,x2) = 0x1 + x2 + 0
i.e. C1 = 0; C2 = 1; C3 = 0
suffices, where a result v > 0.5 indicates A and v <= 0.5 indicates B.