Datasight.io – Machine Learning for the Masses
datasight.io
datasight.io
More importantly, any valid ML analysis requires that the source data is good, which is impossible to guarantee with this service.
I've found most people who need machine learning usually want control of it/hosting it in house.
We'll also deploy to an HTTP endpoint and you can call it and get JSON response.
Does that address your concerns?
I know you guys are only targeting the very simple needs of people which is great. Keep in mind I'm far from your target audience being almost a little too deep in the machine learning side. Best of luck with it!
What value add would you want to see?
I've found that certain things, especially NLP pipelines, are usually very tricky to get right.
Let me give you one example from my day to day:
Sentence segmentation/tokenization ---> part of speech tagger ---> filter words by part of speech and then rank by tfidf scores for topic relevance scoring.
Computer vision also has its own problems ---> binarize images for detection of certain kinds of features vs needing colors, so different kinds of image transformations for handling of different kinds of object recognition, or scene detection.
I also do churn prediction, by the time you're done vectorizing users, why should I take time out of my day to then upload all of this to an external service when I could just run logistic regression, random forest or what have you locally? I typically calculate profit curves and the like as well.
Re: your word vectors and deep learning. I do this in my distributed deep learning lib I hand wrote myself[1]. In my word2vec computations, I have found it takes a significant amount of data to tune right, I usually want control over this pipeline as well.
Again, not your target audience ;).
To give you some helpful feed back, don't try being everything to everybody, own a certain niche really well and run with it. If you'd like to take this discussion offline, I'd be more than happy to help, email is in my profile. Again, good luck with the service!
Let's say you've got 2 steps: Prep the data, build the algo + infrastructure. Let's say for the sake of argument that they're about equal in terms of difficulty. If you can drastically reduce the amount of time the second takes, that will still be a valuable speedup. And I'll bet there's tools that could be included in these types of solutions that can help with data prep also (it doesn't have to be one or the other)
Databases for instance -- there's lots of options, some proprietary, some open source -- but you'd rarely consider building your own unless as an intellectual exercise. At what point does this become true for machine learning? Saying it doesn't do 100% of my job for me doesn't mean that it won't work as a product category.
Of course, this is coming from someone with a vested interest in the category, so take that with the requisite bias...
Still, we've constantly been surprised by how well we can learn on raw CSV dumps from SQL joins or just log data.
Furthermore, small business owners or non technical folks can still understand the concept of "teaching" an algorithm: things like including demographic information increase predictive power etc.
I'd love to get others thoughts on this and if they would use a service like this what you would want in something like this.
I would imagine something like churn prediction as a service, face recognition as a service, sentiment analysis,... more concrete needs could be good.
The landscape is mainly APIs (wise.io,algorithms.io,..)
BigML does decision trees as a service which at least allows you to interpret the results with a good visualization.
Ersatz does neural nets as a service with a decent GUI and an API.
However, there's nothing at this page right now that would convince me about this service. There's no actual examples of what your technology can achieve, there's no programming API, just a promise. I understand that you're still in development, and I'm not asking for full documentation and a page of famous logos under "they use us", but at least a code snippet or something like that would make it easier to visualize.
Right now, I'm lost; my data is a lot more complicated and inter-connected then a mere spreadsheet, will you support it? How will it work? What pricing will be, and how will it be calculated? Yes, I understand that you haven't decided yet on all this things, but that's the problem: there's nothing to show except for the idea yet.
Anyway, I really hope that this service will work great, and I wish you all the luck with it.
The flow is simple - upload a CSV and we'll spit out serialized models, CSV of results comparing algorithms, and an HTTP endpoint where you can get predictions.
To your point about being more than a flat CSV: yes that is absolutely correct. We're starting with flat but moving towards sucking in relations (automatic joins) and constructing feature vectors from that.
It's still in beta, but it's a lot more full featured than it was a few months ago. We're trying to make deep learning easy and we provide a variety of neural network algorithms backed by GPUs. Documentation still sucks... I'm co-founder.
It's a tricky balance w/ all of our products right now (BigML, Ersatz, prediction.io, datasight.io, wise.io, more) -- how to strike the right balance between appealing to Excel power users versus people who are very comfortable in matlab or python but still don't want to build their own if there's a significantly easier option.