Ersatz - Deep neural networks in the cloud
ersatz1.com
ersatz1.com
Why? The real pain in the ass in training a deep network is the hyperparameter selection.
What is your learning rate? What is your noise level? What is your regularization parameter?
Choosing these values is a far bigger pain than almost everything else combined.
Doing a grid search is intractable. Random hyperparameter search is better. You can use a sophisticated strategy, like Bergstra et al have proposed.
I haven't played with automatic parameter selection much (but have been seeing more papers on it recently) so I hadn't really considered it all that closely.
While I'd like to give people a fair amount of control over model parameters if they want, it probably is very important that I make things as turnkey as I can. Shouldn't be too tough to hack something together and make it an option during training.
While I'm trying to start things off relatively simply, the overall goal really is towards allowing people to create models that act as parts of much larger systems, maybe larger neural nets themselves. A sort of genetic algorithm that spawns new neural networks with random parameters and random connections to previous networks could be kind of neat, and making the base elements of those types of architectures (a single fully connected deep net, for example) easily accessible is a first step towards that goal.
see Hinton's google code slides for more info on how powerful these things are:- http://www.youtube.com/watch?v=AyzOUbkUf3M (that's 2007, things are even spicier now)
1) Cloud GPU computation where you upload some special model code that is run on the neural network? ie. your own code
2) Upload data and run some pre-specified models on it, such as in the example you have a '-d model=spanish_speech_recognizer' - in which case the offering is all about how many and how good your pre-defined models are.
The two different use cases are for completely different target audiences.
Parameterizing & fitting these networks (both in having a good underlying representation, and deciding which actual parameterization to use) gets tricky & requires some domain knowledge when you start doing things like conditional or continuous RBMs.
So basically, you bring the data, pick the neural network architecture you want to use, and set its parameters. The model trains on the data you've given it using a GPU cluster (which still takes a while)
'spanish_speech_recognizer' is the name of the model you just trained, where 'MRNN' is the actual architecture (a multiplicative recurrent neural network as described in http://www.cs.toronto.edu/~ilya/pubs/2011/LANG-RNN.pdf) used in the example.
So the models themselves aren't pre-defined, but the architectures you can use are. You can play with a lot of different parameters (if you want), but you don't have to worry about optimizing the code for GPU or making sure your implementation of the algo itself is correct. At least that's the idea.
Although I should add that the response so far has been way beyond what we thought it would be (which is fantastic!), so it may take some time to get to everyone. The beta literally just opened this morning, and I'm not ready to open up the product before working with beta users to polish it up.
Latest Firefox on Ubuntu 12.04.
On one hand, AWS GPUs do seem a bit slower than bare metal, and it's theoretically (actually?) more expensive. Plus we get more control if we host our own.
But then again, running a data center is a problem that has been solved and I'm not sure we can create much value there. In practice, it will probably end up a bit of both, depending on demand.