Deep learning startup Skymind (YC W16) raises $3M, launches enterprise AI distro
venturebeat.com
venturebeat.com
• Kaggle(www.kaggle.com) is a good start for this - start with “somewhat real” problems • Use higher level tools - Keras(https://keras.io/), otherwise easy to get lost in weeds • Consider having a real world goal - eg: if you’re in real estate figure out how to use a simple CNN (not the latest algorithm) for image search • Depending on need consider integration with hadoop/Spark(http://spark.apache.org/)
* fraud and anomaly detection
* recommender systems
* predictive analytics (churn, forecasting)
* image recognition
With image recognition, we hit 98% accuracy on a recent project. Until a few years ago, that was unheard of, and it's simply not possible with other algorithms, so for many companies, deep neural nets can make a significant difference.
Here are two news stories about work we've done for clients:
Making deep learning accessible on Openstack https://insights.ubuntu.com/2016/04/25/making-deep-learning-...
For Canonical, we built a solution that predicts server breakdowns.
For France Telecom's mobile unit, Orange, we built a fraud detection solution using anomaly detection:
http://www.orangesv.com/blog/orange-deep-learning-work-featu...
Knowing data visualization can also be useful.
Most deep learning companies focus on a particular application.
FWIW I'm actually self taught. I did client projects and learned machine learning on my own.
You could start by branching in to data engineering and understanding how data pipelines work. That's closer to the skills a full stack developer is likely to have.
That's pretty awesome and scary in equal measure..
Here's one we use for demos: https://www.unsw.adfa.edu.au/australian-centre-for-cyber-sec...
I am not only talking about Skymind since this issue arises in many other specialties.
How did you come up with the name? Has anyone seen your name and associated it with SkyNet / The Terminator Franchise?
Umm, is that a misprint?
Yes the deal sizes are mid 6 figures or larger.
We have a larger deal pipeline than that though. A lot more to come :).
Red hat/oracle style on premise (non saas) business model.
We usually target NON computer vison applications like fraud, preventative maintenance in data centers (predicting broken machines) and other mission critical applications.
One example:
http://insights.ubuntu.com/2016/04/25/making-deep-learning-a...
This kind of stuff is a swear word on hacker news but there's actually money in it. Fire away if you have specific questions though :).
In machine learning in production there are 2 phases: training and inference (usage)
In training we have spark docker images where you can run cuda right from spark submit.
In inference mode we sit on top of DC/OS by mesosphere embedding lightbend's (they created scala) micrsoservices technology conductr to scale out automatically on a mesos based cluster: http://www.slideshare.net/agibsonccc/deep-learning-in-produc...
Here is more on our enterprise distribution SKIL: http://www.slideshare.net/agibsonccc/skil-dl4j-in-the-wild-m...
If you're curious where the talent is, I cowrote the flagship oreilly book on deeplearning: http://shop.oreilly.com/product/0636920035343.do
We also employ deep learning phds doing everything from deep learning research in health care, ex nvidia, ex cloudera among others.
I'm assuming if Grail (http://www.grailbio.com/) who launched in 2016 with $100,000,000 in funding were knocking at your door you would be more than happy to work with them?
For anyone else we have a very active open source community: https://gitter.im/deeplearning4j/deeplearning4j
Many companies with infrastructure products like ours tend to "incubate" inside a big company first. We chose not to do that. So we spent much of our time just growing the user base first.
From reading about the company it seems that you want to monetize on ease of use with the distro? Do you have customers that specifically pick you because deeplearning4j is open or do you find that it's more of a nice to have or even don't care as long as it works type of situation?
I'd also love to hear some reasoning for picking Apache 2.0 (i.e. a non-copyleft one). I've talked to at least one FLOSS founder who would have picked a different license in retrospect but feel like mostly the license matters less than most people think.