Machine Learning Is a Marvelously Executed Scam
lastweekinaws.com
lastweekinaws.com
But when ML works--Uber forecasting demand, Instacart optimizing shopping paths, delivery service route dispatching, and an infinite number of other real-world use cases with double-digit process improvements--it's extraordinary.
Sales forecasting, fraud prediction, modern search, and more are all powered by ML, and throwing a few stones at bad business plans, bad marketing, or over-marketing doesn't change the fact that those things are bad precisely because they're not that good at actual machine learning.
Certainly, the number of times I've seen business stakeholders treat ML as magic, where they say "It's like (standard business process) but we'll use ML!" to try and create a business case is appalling. And I think that's more what he's referring to; in many companies, ML is a solution in search of a problem, one that the business is quite happy to pay for to say they're doing (it pleases stockholders), and data scientists are happy to accept money for (it's a job, after all).
Classical algorithms for path-finding for example might work really well in narrow cases that have firm constraints. ML allows you to expand the scale of optimizations arbitrarily.
It's kind of boring to read "our revenue jumped 4% after adding multi-objective optimisation to our existing model", but if you stick a few 4% improvements together and apply them to a big revenue stream, you get a big number.
We are an industry as notorious for snake oil as we are for our lack of standards. And both have standard apologetics. The parent's comment is reminiscent of the refrains "Agile, you're doing it wrong" and "that's not the way to write microservices," among many similar others.
I don't think the problem with any of this is ML as such. ML is just software. But software has problems.
A govt. department contracted out the development of an ML model to identify an invasive tree species from sat and aerial imagery. New Zealand is sparsely populated and mountainous, so crews are deployed for a week at a time by helicopter to remove these trees. It is very expensive. By being able to scan large amounts of the country for these trees, they can optimise the removal. The model appears to work very well and can identify the trees when they are young.
These kind of deep models are hard to do with traditional computer vision.
How do you arrive at this conclusion? What about Computer Vision? Speech recognition and synthesis? Fraud detection? Sentiment analysis?
However, AI/ML is definitely opening up new areas of business, such as autonomous robots (A roomba is much better at its job with a ML component in it, without which, it would have been difficult to come up with such a device), Social media, face recognition, etc.
All in all, most of AI/ML needs are centered around pattern matching in one form or the other. There is much less AI, a lot more Pattern Matching.
ML seems to be something an organization reaches for most often when: a) it doesn't understand the data it has and doesn't have individuals who have the competencies to hire analysts to help them; b) when they want to tell a story to an audience of engineers (e.g. as to attract "talent" or signal about how bleeding edge the tech is in the company).
A far less frequent case is when individuals with actual expertise have identified a real need for the use of the statistical methods and infrastructure used in ML applications. In this sense it's a scam--but one that technology organizations use against themselves.
The real money in ML, just like with most fads, is in selling the tools, not doing the work. Hence you see the rise of all these "ML ops" platform type businesses or business units (see: Sagemaker, Databricks, and various others).
Yeah, this is common and often explains why ML projects fail. ML won't magically understand data for you. And if you don't understand the data you are feeding into your ML pipeline, you will almost certainly have a garbage-in-garbage-out situation on your hands.
Hiring a data scientist before you have a solid data engineering pipeline is like hiring an interior decorator while you're still framing your house. Unfortunately most businesses (even highly technical ones) just don't understand the moving parts of analytics.
As an example, I recently discovered a bug in a production system that was costing many millions of dollars, that essentially happened because the team was told to go off and implement a shiny new ML model rather than understand and incrementally improve the system.
It's incredibly depressing.
I think unexplainable decisions affecting the lifes of of human beings (e.g. Googles Playstore App removal algorithm) are today's version of the "Terminator" movies.
I mean.. Yeah..
I'll just say it's unimpressive low hanging fruit to write an article about any wildly popular subject to rehash basic mental models-as-objections-to-hype like Sturgeon's Law, Maslow's hammer, Occam's Razor, and YAGNI.
ML is objectively great at finding signal in noise. When you know what signal you're looking for and how much it's worth, ML is very economical - at certain scales, it's the only sustainable option.
Here are 6 steps to evaluating ML success. The most important point is after each step, if you don't have the economics right, stop what you're doing and go use your money for something else.
1) Choose a pattern recognition objective with a financial benefit Quantify your objective over a fixed timeframe. Focus on how the objective will be achieved. "If we can reduce MTTR by 35% and reduce our fixed maintenance cycle time by 15%, we will see $11.7m gains in net revenue over the next 5 years by reducing our maintenance crew contracts."
2) Draw a clear picture of the data you have and its relationship to the objective. Identify the target variable(s). Assess the quality and granularity of your labeled data. Identify relationships between your data. Declare a hypothesis on the marginal relationship between the model loss function and the objective. Eg: "If this model is 90% accurate we expect to see a 6% lift in upselling / increase maintenance cycles by 14%, etc."
3) Bring a devil's advocate into the room and let them try to tear the idea apart. See if you can piece it back together. Compare your theoretical signal to whatever heuristics you have today. Why is it different? Can you just automate those?
4) Do a 6 week pilot to find the signal in the noise to prove your hypothesis. Do not worry about "cloud scale data platforms." Do not do MLOps. Do not build pipelines. Move heaven and earth to get data out of source systems quickly, by hand if you need to - copy and paste is also data engineering, have a domain SME and a data scientist joined at the hip for avoiding rabbit holes. Use a Kaggle-style holdout set for final model evaluation.
5) Do a field pilot, with your devil's advocate as the judge. A/B testing is usually best - think John Henry vs. the steam engine. Again, move heaven and earth to test your hypothesis in a real world setting. Bathe in uncertainty and confidence.
6) Estimate the cost to create and operate the ML pipeline needed against your expected benefits. Is it justified? Build it; otherwise, either look for ways to augment its value or kill it.
Most ML goes astray at steps 1 or 2. But a lot of good ML solutions are missed because the "elegant scam" skips most of these steps altogether.
Just the rekognition service he is whining about can and is being used to detect cheaters in carpool lanes and automatic billing. Postal mails are sorted by ML, as well as spam at an email service provider. Recommendations and search on an ecommerce shopping site directly multiplie the revenue of the site. I had no idea some one who calls himself a software engineer could be so completely clueless.
I expect this person to delete this post in the next couple of years when he realizes how dumb it is.