Explaining machine learning pitfalls to managers (2019)
gudok.xyz
gudok.xyz
It goes something like this: someone (a PM, manager, or sometimes an engineer) has a brilliant idea to use ML to enhance some part of the business, say automate a manual process. A random folk is asked to do a PoC, and they slap together a model in a couple of days.
This model often shows impressive performance, say 80% accuracy in a problem where 90% is considered acceptable. Leadership gets all excited and they sign off the project. And they get themselves into a world of pain.
The pain takes many forms, but the most common ones are:
a) That extra 10% accuracy is extremely hard to achieve.
b) Running ML in production is really difficult and the company doesn't have the skills/maturity/expertise to run this new, possibly mission critical components.
The pitfall that lead to this situation are:
1) Assuming ML is easy from a PoC.
If we built a PoC that achieves 80% accuracy in 2 days, how hard can it be to achieve 90%? it turns out it can be really difficult. Performance improvement of ML models is not linear and it can be really difficult to get even a few % points better.
The second part is running ML in production. It might be obvious to an engineer that there is a big difference between slapping together a prototype in a hurry and running a mission critical service in production, but people unfamiliar with ML tend to assume that the process of building the model is all there is to it.
2) Assuming ML is not all or nothing (for your particular problem).
One might tend to think that a model with 80% accuracy is just a little bit worse than one with 90% accuracy. However, depending on the domain, this isn't true. For some problems models need to perform better than a certain threshold to be of any use. In that case, a model with 80% accuracy is as good as one with 0%.
This is an oversimplified explanation but I've seen it happen often enough that I consider it an (anti) pattern.
The worst is when management just can't let go of an idea and it becomes their white whale. The project I'm thinking of has gone on for 5 years now, been worked on by consultants and a few teams of data scientists that have come in through acquisitions and all quit. The only thing that's stayed constant is the shitty data, the dead horse and the management beating it hoping to get a better AUC.
- The humans could handle about 1000 events per day across the team.
- The system received around 100 million events per day.
- This meant that something like a 99.999% (5-nines!) true positive rate was the target goal.
But management poured millions into the team over the years because somebody had come to believe that the machines could handle 100% of the challenge and had gone to some AI conferences and become a true believer.
The correct approach was team the analysts and the machines together in more effective ways (triage, enrichment, context providers, etc), and the correct goal was to focus on improving the results of the joint system over just the human team, not reach for absolute perfection. Triage, and accepting that sometimes shit gets through and use good practices like layered defense, n-factor auth, and zero-trust systems would carry the day.
I remember the presentation where the new head of the cyber team explained that it's acceptable for trojans that target Windows 98 to get through as they had no vulnerable systems the trojan could exploit -- the entire company was MacOS and Linux. You would have thought heads would have fallen off of the executives.
Going from 80->90 is cutting the error rate in half. That’s quite a jump.
This is where I often get to when thinking about ML or AI (or AGI).
We do so much training to make systems predictive. I think it is unlikely that we can ever train it so much that it it can become predictive on its own.
As said in other comments, going from 80 to 90% accuracy is a big step. 90 to 95%? Thats too much to ask.
I'm human and I fuck up way more than 5%.
Though I may disagree with his classification hard / very hard / very very hard: you can definitely work on some ranking tasks with just behavioural data (i.e. no labels). And to build decent OCR, you need to label your own datasets, the problems is not that different from speech.
Also, what is hard / not hard changes w/ state of the art. Translation is a good example: it used to be that you had translate 1:1 each text. But not anymore w/ recent advanced in NLP (transformers, etc.).
Great job!