Some Starting Points for Deep Learning and RNNs
aistartups.org
aistartups.org
It feels like most examples of commercial AI usage are features (image search, automatic face tagging, etc.) that get added to existing products by large companies with access to large datasets and computing power.
Some big deep learning projects might be beyond small projects/startups, but I think AI is reasonable if you work methodically.
edit: $50, not $100.
[1] https://medium.com/@samhavens/building-somerset-d518ba284c49...
Took a look at the site and I signed up, some words of caution for anyone looking for the paid edition, but it works well for me:
* You cannot display this information publicly in the free tier, and you must display their logo for attribution on their page at a specific size they mention in their TOS. * They want your phone number for some reason.
* On the other hand, the api (AYLIEN https://developer.aylien.com ) either changed their limits or you might be just mistaken, it looks like the free tier is 1000 calls a day now, which is awesome!
Good introduction to the field: http://ocw.mit.edu/courses/electrical-engineering-and-comput...
I'm going to make the simplifying assumption of saying that "machine learning" is a subset of AI for our purposes, and that "machine learning" includes "simple" techniques like linear regression, logistic regression, k-means clustering, etc. I don't think any of that is terribly controversial, although I know a few people would quibble over it. Anyway...
Given that, I'd say the example is absolutely "yes". AI can help startups and small projects. There's actually a really nice example that Andrew Ng talks about in his Machine Learning course on Coursera. In that example, a t-shirt manufacturer is trying to figure out how to size their shirts. So they go out and take a bunch of height/weight measurements of prospective customers, and then use k-means to segment the data into clusters. Let's say they do 5 clusters, corresponding to extra small, small, medium, large, and extra large. Now they can look at the measurements in each cluster and figure out how to size their shirts.
More generally, anywhere that you have data, and you want to extract insights from that data, you likely have an application for some level of machine learning / AI. Of course it won't always (or even often, perhaps) be the case that you need a deep neural network, or anything like the cutting edge in academic research. But, then again, you might. :-)
that get added to existing products by large companies with access to large datasets and computing power.
Do you have an AWS account? If so, you have access to large datasets and computing power as well. Look at all the data that's "out there" in terms of Open Data, LinkedData, etc... As an exercise, browse around data.gov sometime, and look at things like the datasets the UN makes available, and the world bank data, etc., and see if you can think of a way to combine a few of those datasets and extract some meaningful insight from a combination nobody else has looked at before.
If you come up with something, it's not terribly hard/expensive to spin up a cluster using EC2 and run some analysis.
We do research consulting in speech & language, and tons of small start-ups come to us with a variety of really interesting problems.
cobaltspeech.com
I'd expect trained classification to provide automated curation to be the best way for startups to take advantage in an existing project.
Even for people that do want transcription, though, often you can handily beat the accuracy of off-the-shelf services by tuning the models to the domain.
I'm pasting some copy from our website in part because I can't talk too much about specifics, but here's some examples:
"
High quality speech to text transcription, including in the very difficult areas of conversational human-to-human speech: phone calls, voicemails, meetings, etc. Classification of speech & language: identifying gender, age, regional accents, level of education, etc. Voice front-ends for various applications and devices, including natural language interfaces: heads-up displays, robots, smartphone apps, etc. Speech synthesis (a.k.a. TTS) to generate high-quality synthesized speech from text. Analysis of audio: detecting different types of noise or speech: dog bark, shouting, gunshot, water running, cars, etc. Customization of speech recognition to processed signals, such as proprietary compression or noise reduction algorithms Fine-grained analysis of speech and language, for giving feedback in language pathology or language learning. Analysis of speech & language to detect various health conditions: stroke, dementia, depression, schizophrenia, etc. Information extraction from text or audio: phone numbers, dates, entities and relationships, etc. " [0]
I'm trying to be vague because I don't want to violate NDAs and it's a pain to think up examples I'm not currently working on, but there are a lot of applications of machine learning to the _speech signal_ (rather than to the problem of speech to text; there are many other things you can predict) that are super interesting. Some examples are listed in the text I copied.
> 'a half-baked siri' There are so many instances where speech-to-text is super useful for applications other than consumer facing CRUD apps. In a lot of varied industrial applications people are doing very cool things with speech; there are a lot of applications that need very targeted speech systems.
Disclaimer: temporary contractor for one of them.
So far, I've only really delved into NNs with one hidden layer, and I read that problems really only require one hidden layer. For what kinds of problems are multiple layers required, or CNNs, or RNNs? It seems that there's a lot of "cool factor" there, but I don't know what they bring to the table that's new.
Where I read about one-layer being "good enough" for most problems: http://stats.stackexchange.com/questions/181/how-to-choose-t...
If you work with already extracted features (like a classical ML pipeline where you can use SVMs or logistic regression as well), you may go with one-layer network. But deep learning really shines with perception-like problem.
Don't be fooled by the approximation theorem (you can approximate any function with only one hidden layer). It is beautiful in theory, but in practice it is like saying that you can write any program in brainfuck because it is turing-complete.
1. The pace of published research is pretty fast right now. This makes it difficult to know where the research fits in when solving problems. It'll probably take a few years before we know where to use many of the approaches published last year.
2. Iteration performance (trying many new things quickly) is improving with high-level frameworks but still lots of work to do here. Since we don't know where a lot of research fits it's not always apparent which methods work best (and "best" changes every quarter).
3. We're still missing theoretical foundations for much of deep learning. This is useful, not just for research, but to know what can and cannot work with current approaches.
4. Model architectures are still based largely on trial-and-error, intuition and search.
Some hard long term challenges revolve around cases where you don't have a lot of unlabeled data or examples of classes you care about. There's also the technical challenges of training large scale models without obscene computational resources.
I also prefer to reserve the term AI for generalized AI, which we're still a ways off of, as opposed to modern classification problems etc. that I would call machine learning (though I know that nomenclature is uncommon).
EDIT: jimfleming makes a great point about theory as well - we could likely be much more efficient with better theory for deep neural nets.
That said there is a bigger problem with those bounds because it doesn't incorporate model complexity. The VC dimension is much more insightful because the complexity of your model and the hypothesis space it represents is important for proper training. As an example, add a regularization term to your model and you're no longer doing anything like N/e^D. Convolutions, dropout, etc all prevent NN models from becoming too complex to train.
People often try to intuit the intrinsic dimensionality of a dataset by using techniques like looking at singular values above some threshold or reconstruction error versus changing the output dimensionality of a dimensionality reduction/unsupervised technique like PCA, matrix factorization, or an autoencoder.
An info theory person might argue entropy and compression ratios are also insightful.
Speech was one of the places where HMMs were _huge_, so that's kind of a big deal.
[0] http://googleresearch.blogspot.com/2015/09/google-voice-sear...
HMMs also used to produce more meaningful intermediate results, although I guess these 'DeepDream' images set a new record in that regard :)