Microsoft launches a drag-and-drop machine learning tool
techcrunch.com
techcrunch.com
ETL-style dataflow pipelines are more natural for such tasks than imperative programming, it's not just "for people who can't code". Rapidminer is dataflow-based too.
Actual links, if you want to avoid gadget news website:
https://docs.microsoft.com/en-us/azure/machine-learning/serv...
https://docs.microsoft.com/en-us/azure/machine-learning/serv...
In general the holy grail is an autonomous ML system. with no input from the user but a data set schema.
> This tool, the Azure Machine Learning visual interface, looks suspiciously like the existing Azure ML Studio, Microsoft’s first stab at building a visual machine learning tool. Indeed, the two services look identical. The company never really pushed this service, though, and almost seemed to have forgotten about it despite the fact that it always seemed like a really useful tool for getting started with machine learning.
> Microsoft says this new version combines the best of Azure ML Studio with the Azure Machine Learning service. In practice, this means that while the interface is almost identical, the Azure Machine Learning visual interface extends what was possible with ML Studio by running on top of the Azure Machine Learning service and adding that services’ security, deployment and life cycle management capabilities.
The answer is "not much."
It always turns out that the difficulty of "setting up a basic [hello world application]" is entirely unrelated to the essential complexity of the problem space and attracting a broader range of new users is later viewed as a valuable advance.
A fluent Python developer that doesn't understand basic ML concepts can easily use something like Scikit to code and build "wrong" models. By "basic concepts" I mean standard tasks like data preparation for specific algorithms, sampling, evaluation methods, testing for bias, or just generally how to properly execute most ML tasks like the ones prescribed by something like CRISP-DM.
Someone with basic coding skills - e.g. knows SQL and some imperative programming - but with a solid understanding of ML tasks, and how to execute them properly, probably has a better chance of coming up with better results than the former, using something like IBM Modeller or RapidMiner.
Note that I'm not saying that a drag and drop tool is superior; you could build a flow-based GUI for Scikit, so a tool like this is always, at most, an interface to some code libraries (Scikit, in this example). Having full access to the actual lib, or just better libs, is likely to be less constraining, and more apt for more sophisticated approaches.
But beyond that, I’m going to use the grown up tools to build real things (and get paid for doing so).
I'm a fan of plain text and source control myself.
If not, it looks like they have parallel projects working on basically the same thing. I’ve been waiting for Microsoft to open Lobe to the public..
1. They want to build, train, and deploy a machine learning model into production. Presumably as a microservice, part of a web application, etc
2. They don't know how to program
I honestly can't imagine a less useful product than drag and drop ML.
I do know how to program.
I still want this tool. Badly. So badly. Just because I can program doesn't mean I want to use programming to solve every problem. I want simple tools that do the complex things for me for the 90% of cases where they are good enough.
Let me spend my time on the 10% of problems that simple tools can't solve!
It seems like this would be the worst of both worlds. Too simple to be of any use to any ML engineer, too complicated for the uninitiated, not customizable enough for a domain expert.
I haven't used this tool, but it seems reasonable that the people who it might be useful for (regardless if Microsoft PR recognizes this or not) are for scientists who have an idea for an algorithm, but don't want to spend too much time thinking about how to write and deploy python/C++ code to their clusters.
Sure, it may not seem like that much effort to many programmers, but as the Fortran discussion stressed, just because you can code and think logically, doesn't mean you're a programmer.
I’m not interested in deploying a service or anything, but it would be a great way to take a first pass at analyzing some of the huge and pretty complex datasets that we generate, like metagenomic DNA sequences of microbiotas that are paired with health related information that could also be fed into the model.
Even just narrowing down a list of potential targets would be pretty darn useful.
Not a lot of these people can program.
I know a lot of programmers who want to work with ML, but in my experience, very few programmers are good enough at math or statistics to do so, and even fewer have the business skills to actually translate their results to management in non-tech based organisations.
I’m sure a lot of programmers will make excellent data-scientists, but I’m not entirely convinced why I would bring ML to my programmers rather than my people who have degrees in applied statistics and organisations.
I work in the public sector. One of the reasons ML hasn’t found it’s golden case yet, is largely because no one have figured out how to use ML in a way that is better than the decades worth of data-related work we have already done, and part of the reason behind this, is that companies who sell ML are programmers. They know how to use ML to identify, but none of them, not even IBM seem to know how to use ML for something they can actually get us to buy. And lord knows both sides of the tables have tried, I mean, even our political leadership has heard of the ML hype, and want us to use it. So I’m rather hopeful these tools for non-programmers will bring ML into the hands of people who will know what to use it for.
"Anyone smart enough to use this isn't dumb enough to need it."
Those were the words that come to my mind everytime I see stuff like this. Clickers want solutions, not complicated toolkits in a GUI.
I don't see what the market would be for implementing this into production, like you say to make live in production needs to know how to code. Perhaps though I could see Segment offer a similar tool which I guess would be 'in production' without code.
I still use their free Jupyter Notebooks service also.
Maybe Microsoft aims at teaching ML to beginners, which still would be detrimental if they get used to just that.
In my opinion, it's far too easy to make a critical mistake during design/implementation of ML to follow this same path. And what's more, if you mess up making an analytics dashboard, it's usually fairly obvious. In ML, there are MANY ways to mess up a model and you have no easy way to tell.
If someone doesn't have the technical experience behind creating these models, I would not trust any output they give me from using one of these tools. And if they do have the experience, they would certainly not be choosing to use one of these tools either.
I am building a competing tool, so I am not affiliate with MS, but I do think that auto ML has value.
Machine learning is different from imperative programming in such that most of the "programming" is done by experiments and not with actual "program", hence there is an opportunity to replace programming with compute. I.e. an automl platform can create 100's of models/pipelines and just try them all.
Also, why would you trust a model which was created manually and not a model which was auto created.
When a model is created in auto ML it pass the same validation process as manually created model, so in both cases the quality of the model should be judged independent from the way that it was created.
In addition, all models (regardless of how they were created - human / not human), should be monitored for predictive performance. I.e. I will not "trust" any model without continuous verification.
There's no question that there's value in AutoML system yet most ML production systems I've worked on / seen were way more complex than feature vector -> model -> prediction. You likely have multiple models, pipelines, normalizations and plain old conditionals. Hard to automate all of this.
Note that automation is not only building the model, but automating the full life cycle - pre processing, hp optimization , pipeline deployment and monitoring/retraining.
the short answer is, go study stats and fundamentals of ML instead of asking hn to build your product for you.
> "why would you trust a model which was created manually and not a model which was auto created."
one of many reasons: domain knowledge is important, and math alone cant tell you things are muffed up. contrived example: you build a linear regression model to predict home price and square footage has a negative coefficient. Math conclusion: bigger house = lower price. domain knowledge: oh, we are missing a feature and the model cant tell the difference between city homes vs rural.
there is value to auto ml but there is a lot of room to go horribly wrong
You are pointing to an area outside the realm of automl (feature engineering/generation) , which is domain specific. But this was not my original question.
You could argue (which you didn't) that this would fall under model interpretation, but a model in this example would probably fail to generalize and make bad predictions in the future: IE slamming home values because they have large square footage.
What about all those businesspeople who only hire analysts to tell them (and their peers) what they want to hear? Now they can tell themselves what they want to hear, having laundered it through a computer.
I've seen many people in ML not quite understanding what they are doing randomly trying different things until it "works".
The problem gets worse in unsupervised ML, e.g. cluster analysis. Whatever variables you choose, clustering will give you some results. But only an experienced person can understand what variables to choose for the clustering, how to do it, and what those clusters really mean. You can't just try different things in clustering until it "works", because it always works.
Really? How many of the machine learning "specialists" actually write custom codes for matrix operations done by the ML engines?
B-t.
> at least those who don't just copy-paste code snippets from SO
This is like the late 1990's listening to "real coders" complaining about web developers using Javascript all over again.
https://docs.microsoft.com/en-us/azure/machine-learning/serv...
https://en.wikipedia.org/wiki/SPSS_Modeler
I remember using this in NINETEEN NINETY FIVE (it was called clementine then).
Interestingly one of the leads (Rob Milne) sold up (to IBM , forced sale I guess due to a cash squeeze and no investors) and went to Everest, where he got to the bottom of the Hilary step, had a massive heart attack and died.
Makes ya thunk.