How to deliver on Machine Learning projects
blog.insightdatascience.com
blog.insightdatascience.com
Now our company doesn't have any machine learning expert or a data science genius. Going for hiring one would take time. Taking someone up on contract would be very expensive (our CEO wasn't ready to shell out that kinda money). So the task fell on me. They asked me to go through the multitudes of Machine leaning MOOCs out there and get a working prototype ready in 2 weeks.
I had already done Andrew Ng's course back when it came out for the first time. But my memory had faded for the lack of practice.
I re-ran the course again. I went over a couple of online ML books too.
Then I started thinking of the problem at hand. Unfortunately, it turned out to be a chicken and egg problem. For the feature to work perfectly we needed a large amount of training data to train our models. But without the feature actually deployed, we didn't have any way to collect any training data.
So we ultimately fell back to simple algo, that took it's decisions based on a few hard coded rules. Things have been working fine till now.
In my experience, having someone that knows what they’re doing on the front end of a study design wise can save weeks or months of work on the back end of a study or project.
Oh c'mon. Any large company today and the expectation or deadline for practically anything is "asap" or measured in a few weeks at most. Short-term thinking is a major player in publicly traded companies. Because of that, this is what opens the door for startups to play the long-game.
Since no one of you had any experience with ML, how did you know that a ML algo (which one?), implemented "somehow" would give you the results you wanted? (Not a cynical comment; I am really interested in hearing about this).
That said, there's a good chance that your current algorithm is all you will ever need - many times a ML project is too much, and you already have good results.
Designing the perfect viola using machine learning doesn't sound like it's something for beginners.
"Viola" either refers to a stringed instrument, or means "raped" in the sense "he raped" ("il viola"). So please don't use it as an interjection.
In addition, trying for the feature to “work perfectly” from the get go, even with lots of data usually is quite hard.
In many ways, traditional approaches were harder because you need huge amount of domain expertise in CV & NLP, whereas a ML expert can solve simple CV problems with almost no domain knowledge.
Now, a lot of the business data, especially time series data, I agree that an algorithm/heuristic approach is easier and more robust. E.g. recommendation systems.
Not sure what the parent meant by "algorithmic approach" though.
Everyone outside of data science seems really surprised by this and I can't count the number of times someone has asked me to build an algorithm for X but has none of the data to support doing so. It doesn't mean the feature/product can't be built but they often want a supervised learning solution without the cost (and time) of acquiring the ground truth data.
Neural nets are basically black box heuristics, with unpredictable edge cases. Much like human reasoning, I'd warrant!
That doesn't mean it's bad.
Many resources exist online about how to get a model to converge, and that’s not usually what makes or break a project.
Data acquisition, augmentation, model selection, and iterative exploration however seem quite rarely discussed compared to how important we have seen them be. This is our attempt at sharing this outside of our usual circles.
https://en.wikipedia.org/wiki/DMAIC
Nothing wrong with that though...
Seriously - an optimisation loop on a test set? Seriously?
Edit: typo "of" to "if". Somewhat serendipitous if you think about it.
It's just funny that "Data Scientist" seemed to be originally branded as the more technical/engineer-y version of a data analyst. Now I get recruiters contacting me for "Data Scientist" positions that entirely revolves around SQL and excel, and nobody in the Bay Area hires "Data Analysts" anymore.
Alright, guess it's time to update my LinkedIn and resume to adjust for this inflation? Maybe I should jump up a few inflation levels and just become a "Deep Learning Engineer."
Taking advantage as much as possible of hypes and other people's lazyness is fine in my book. It is certainly not my duty from the outside to educate recruiters and business people who make hiring decisions on the field – when I tried, from the inside, to gently point out that what they were thinking did not make any sense, I just put myself in a dangerous spot. I can be a data scientist, deep learning engineer, machine learning engineer, machine learning research scientist, whatever pays more and whoever has the most fun. If using an RNN instead of a more effective and efficient linear regression gives me more money and prestige, I will do it – as an IC you either go with the flow or you are not having a good time. The vast majority of us is not saving lives anyway.