The First Rule of Machine Learning: Start Without Machine Learning
eugeneyan.com
eugeneyan.com
"When you have a problem, build two solutions - a deep Bayesian transformer running on multicloud Kubernetes and a SQL query built on a stack of egregiously oversimplifying assumptions. Put one on your resume, the other in production. Everyone goes home happy."
First they marketed it heavily before even thinking. During test cycle they fed the entire data corpus in and ran some of the original test cases and found some business destroying results pop out. The entire system ended up a verbatim port of the VB6 crap which was a verbatim port of the original AS400 crap that actually worked.
The marketing to this day says it’s ML based and everyone buys into the hype. It’s not. It was a complete failure. But the original system has 30 years of human experience codified in it.
So yeah anything that's still running on VB6 is very likely crap.
At any rate even when it was maintained still looked down on, I guess a Dijkstra based side-effect.
No, it does not. It means that Microsoft no longer provides support for the IDE. That does not prevent the developer from maintaining their own VB6 code. With some extra steps, the official IDE and compiler for VB6 can still be installed on Windows 10. Running programs built from VB6 is still supported.
> written 2 decades ago with the coding standards of the era, the original developer team is gone and probably retired
This applies regardless of the programming language to any codebase that has been around for long enough.
If you're comfortable with this then I don't think you're actively investing in your software.
>This applies regardless of the programming language to any codebase that has been around for long enough.
No, if you have a team actively maintaining the project you have the knowledge transfer in-house which is the second part of that sentence.
What exactly does 'actively investing' mean in this context and why is it needed? If the software is actively maintained so that it continues to meet business requirements, is that not enough?
> No, if you have a team actively maintaining the project you have the knowledge transfer in-house which is the second part of that sentence.
That is orthogonal to what programming language is being used. When the project is actively maintained, knowledge can be transferred regardless of the programming language.
If you're actually investing in maintaining something that's running on a deprecated platform that's decade over EOL and nobody wants to touch with a 10 foot pole - that sounds like a crap project by definition.
Anything that's sufficiently funded to be actively developed would have figured out a migration plan by now, the only scenarios where it wouldn't sound like terrible projects to work on.
>That is orthogonal to what programming language is being used. When the project is actively maintained, knowledge can be transferred regardless of the programming language.
No it's not when the language is deprecated by the owners for over 12 years at this point. It's like having software that only works on windows xp and maintaining it because you can still boot a VM to run it. Good luck working on that POS.
The platform it runs on is Windows 10, which is not deprecated. Microsoft provides an 'It Just Works' guarantee on Windows 10 for VB6 applications. It does not matter whether someone wants to maintain code. The company pays people to do it. Just like how there are many people who do not want to work on proprietary software but do it anyway because their employer pays them to do it.
> Anything that's sufficiently funded to be actively developed would have figured out a migration plan by now, the only scenarios where it wouldn't sound like terrible projects to work on.
Actively developed means that bugs are fixed and features are added as needed by the business. It does not mean jumping on the latest tech trends when there is no business justification. And I am pretty sure that the users are happy that they can use a fast, responsive application instead of a lumbering, bloated Electron app.
> No it's not when the language is deprecated by the owners for over 12 years at this point. It's like having software that only works on windows xp and maintaining it because you can still boot a VM to run it. Good luck working on that POS.
That is a strawman argument because the VB6 IDE and programs compiled with it run on Windows 10 natively, without a VM. And running VB6 programs on Windows 10 is officially supported.
My point is that every time I've seen scenarios like this with products stuck on unsupported platforms is that product is used but there's no money in actively maintaining it (or else it would have made migration plans in the last 12 years). This means you are likely getting shit money working on it, working on a legacy stack you won't use anywhere else, codebase is almost always shit, and the work you do is unrewarding. So crap projects by definition, and every testimonial I've seen so far confirms it.
I've worked on projects being stuck on tech close to EOL - they always had migration plans to upgrade to supported tech.
I have worked on products with old tech stacks as well, e.g. a C++98 codebase for the core product of a multi-billion-dollar company, with no plans to migrate. New features were being frequently added, and the company's own standard library replacement that bridged the gap was itself actively developed.
Even they agreed it was shit though.
We have a COM component written in VB6 running in IIS on windows containers on Amazon in EKS.
It works but it’s crap!
https://en.wikipedia.org/wiki/Visual_Basic_(classic)
>The final release was version 6 in 1998. On April 8, 2008, Microsoft stopped supporting Visual Basic 6.0 IDE. The Microsoft Visual Basic team still maintains compatibility for Visual Basic 6.0 applications through its "It Just Works" program on supported Windows operating systems.
>In 2014, some software developers still preferred Visual Basic 6.0 over its successor, Visual Basic .NET. Visual Basic 6.0 was selected as the most dreaded programming language by respondents of Stack Overflow's annual developer survey in 2016, 2017, and 2018.
Stack Overflow Developer Survey 2016: Most Dreaded: Visual Basic: 79.5%
https://insights.stackoverflow.com/survey/2016#technology-mo...
Stack Overflow Developer Survey 2017: Most Dreaded: Visual Basic 6: 88.3%
https://insights.stackoverflow.com/survey/2017#most-loved-dr...
Stack Overflow Developer Survey 2018: Most Dreaded: Visual Basic 6: 89.9%
https://insights.stackoverflow.com/survey/2018#most-loved-dr...
I disagree that VB.net was dreadful. But it broke backward compatibility but I think for good reasons: arguments being byref by default in VB6, collections being inconsistently 0 based or 1 based, the SET keyword that wasn’t really serving any purpose and was inconsistently applied, having to provide parameters within brackets or between spaces depending on whether the return value is assigned to a variable or not, etc...
I have a lot of sympathy for the frustration of someone who has to maintain a huge code base when backward compatibility is broken, but I think the changes VB.net introduced were necessary.
I remember when my father had to use medicall services billing program supplies by a natinal health insurance company, and he had some problems with it.
Luckily, I was a student in Bucharest and I went to their headquarters to play middleman between my father and their "informatician".
This "informatician" was the sole architect, UX designer, developer, tester, release manager for this program -- VB6+ access.
I sort of helped him debug the code, he built me a special version and handed it to me on a CD.
The program was ok UX wise and blisteringly fast. Years later, they hired this corrupt company that built software for the State and produced a horrendous program, that took terrible and just the startup took 15 minutes (parsing hunonguous XML and inserting it line by line into a local sql database, as far as I remember reading the logs).
The contract ran into HUNDREDS or millions of euros. Granted, the scope of the program was a bit wider.
Look at windows for example. Image how good that would be if they didn’t keep trying to fuck around with it and actually finished something.
Change is needed because people want to have jobs and they will make work for themselves if none exists.
Traditional jobs must have solved this problem somehow. You don't usually see e.g. windowmakers or installers coming up with windows in the shapes of superellipses because square windows are already solved, or stoves coming with integrated fridges because "just an oven and a top" is already solved.
Furthermore, you do see household appliances getting fitted with useless feature bloat and shoddy software and wireless and touchscreens on microwaves etc. It happens. IoT, subscription based software updates for power drills etc... Tractors that can't be repaired and contain a jumble of proprietary software as a service etc.
There was absolutely nothing wrong with the "ugly" framework code, it was quite beautiful, well structured, configurable and fast. Somebody didn't like that you didn't write java code and the properties file based DSL was indeed odd, but nothing wrong with it after you bothered to read the library code.
The Spring Batch code was more explicit, but much uglier, overall.
I'm sure there is far more to the story than the new guys doing a poor job writig a replacement service.
For example, old code does add maintenance costs just by the fact that it's either tied to old frameworks or even OSes, either of which might not be maintained anymore. Also, I don't feel it's a honest argument to criticize java for the sake of being java. If there was ever a production-minded programming language and tech stack, that would undoubtedly be java.
I'm sure the guys who designed the replacement service would have a few choice words regarding the old service.
Well, actually you do. You'd be surprised. Let's ignore for now that there are different window makes and models. There has been a transition from custom-made windows to ready-made windows, whose production cost is higher but the total cost ends up being low because you can have a crew drop by in a construction site and get it done in a few minutes.
Then there's the current progresses in energy efficiency, and also fancy gizmos like actuators and sensors and all kinds of domotics.
Then there's security systems embedded into windows, which further adds sensors and networking and actuators.
You're now far beyond your mom and pop's windows.
And progress ain't done yet. We're starting to see self-dimming windows and also self-cleaning windows.
Find a new need. Every good product (and many bad products) is an answer to some need. And the world's full of all kinds of needs that we can work on.
However, sometimes we start projects without proving they actually answer a need, or sometimes the internal corporate needs don't match the user's needs (I'm looking at you, integrated advertising in Windows 11).
I'm not saying you are wrong, but you aren't right. There is a balance. You can't stand still, but quality that comes from improving the current thing is important as well.
Part of the fascination with ML is the (dangerous) myth that you don't have to wrap your head around a complicated problem anymore, instead the solution will just magically fall out on the other side of the blackbox if you just feed it enough data.
Understand ing the intricates of the problems you are dealing with however is a value in itself.
«Part of the fascination with ML is» solving the mystery behind the ability to automatically build functions and behind those functions.
Surely, both in practice and axiologically, understanding and deterministically solving have a great value. Also because of that, the fact that systems exist that can adapt into solutions, but contain a transparency problem ("yes, but why"), contains an immensely fascinating theoretical challenge, in the learning that may come from the attempt to understand the "grown, spawned" (as if a natural phenomenon) system.
The laziness is not necessary: there is a great deal of fascination in unveiling the mysteries in the blackbox.
Then of course, when you have a practical problem to solve (instead of that intellectual challenge and promise), pick your best solution. And surely it is sensible to call it dangerous to rely on something not properly understood, which may hide the potential faults ("yes, we found out it fails here, and it may be that we kind of assumed it "saw" shapes, while really it "sees" textures..."). In professional practice those "active" fascinations (understanding the spawned) may be luxury.
It's one of the reasons why for specific problems, heuristics or statistics are way better than any attempt at nonlinear modeling / ML prediction (e.g. highly accurate climate models vs. struggling weather models).
We have had a project where we were asked if our model would consider X. So we added X to the model but this didn't increase performance. Now the sane, simple answer would be to just ignore X. But then people come and ask why, doubt that it doesn't improve results, competition without ML considers X.
That doesn't happen (or is hidden) in a none ML situation where some decisions aren't questioned by a benchmark.
But I guess if we complain that half of our colleagues and the media don't understand ML, why should we expect management to?
When the command from C-level is "We need some AI projects to tell our shareholders about," we shouldn't be surprised when middle management suddenly has successful AI projects in their slide decks.
Artificial Intelligence.
Often the biggest benefit is that the ML version is good at catching when the experts hadn't had their cup of coffee as well.
Most experts are like family doctors, they get the correct diagnosis 70% of the time. And even if you juice them up real good, they will ALWAYS lose 5% to human error.
The ML also hits the 70% mark, but it's a different 70%, so it'll fix 70% of the errors. Then you're batting at 0.91 instead of 0.70.
In the end, it gave basically the same results as keyword searching. But we marketed the shit out of it.
If the job involves "looking smart and innovative" for whatever reasons, people tend to err on the side of overly complex solutions.
On the other hand if the advice "let's just go with an SQL query built on a stack of egregiously oversimplifying assumptions" comes from someone, who doesn't know how SQL and linear regression / logistic regression with binning/bucketing / simple decision trees work, I would ask for a second opinion. Because a huge part of the retail banking, non-life insurance and marketing business is running on this simple stack. Obviously profitable.
If the same advice comes from someone, who knows when to use deep learning instead of XGBoost and why, I would go with his/her advice. And I would try to keep him happy and on my team.
The problem was estimating an incoming train speed from an embedded microphone sensor near the train station. The ML scientists used the latest techniques in deep learning to process the acoustic time series. The talk session was two hours long. This project was their showcase.
I guess no one in the prestigious ML team knew about the Doppler shift and its closed form expression. Typically taught in a highschool physics class. A simple formula that you can calculate by hand: no need for a GPU cluster.
The need might be for a sensor local to the platform as a back up to give warning for a train that's traveling too fast? In which case a sensor that mimics the old Cowboy film favourite of putting one's ear to the track seems like a reasonable thing to try.
But this is very likely a very well researched area and there are definitly train people who can point out a flaw in this idea (dirt?)
Particularly strange since moving train (i.e. vehicle) is about the most common way doppler effect is explained in textbooks- it's not like you need any big "eureka" moment to get to this solution either.
Additionally, train engines run as generators to actually power the wheels, which means they're likely running at consistent RPMs or a consistent range of set RPMs. This could be listened for.
In the real world there's often more noise and variance and additionally, part of the benefit of using those techniques is that you can arrive at solutions that are about as good without being an expert in every single thing.
I'm sympathetic as this is a showcase and if their general method performed as well it does show it can learn the data well for other comparable problems without easy solutions. I know I often test my models on verifiable problems as a sanity check..
Also worth discussing is what happens when instead of one you put 30 sensors to improve your estimate of the speed. Good luck figuring out the closed form Doppler expression in that case (you technically _can_ use Kalman filtering but you are assuming each sensor is independent - they would not be, they would be correlated based on their spatial location and closeness to train).
With deep learning, all you need to extend your 1 microphone solution to 30 is a lil bit of pytorch code to add more neurons and some plumbing to pass in 30 audio streams but that's it.
Not to mention extensions to more complicated scenarios - people talking nearby, cars nearby etc. With deep learning you probably wont even need to modify any code, just throw training data (assuming your original model architecture is well designed).
The clues required are in the how the thousands of waveforms are affected by the environment, how they change as the train passes different features, and how their volumes change over time, and other features we can't know in advance. Probably the clicks as the wheels pass joints between tracks are the most telling clues about speed.
The microphone doesn't give a sine wave.
If we can't know in advance, how can you expect a glorified Markov Chain to magically figure it out? If it could - and it can't, but if it could - how would you know it did it correctly?
Fortunately, we know enough about physics to be able to deal with it without a divination server.
I get it. The train operator wants a solution, but realizes figuring this out is too hard, so it's better to pay someone else to do it. That's normal. It used to be that this someone else would do the actual work necessary. But thinking is hard and electricity is cheap, so some figure it's better to just light up a GPU farm and wait until a solution forms in the primordial soup of repurposed vertex shaders. That too, perhaps, would be OK in principle - if the technology was there. But it's not there yet. We're still better off doing the actual thinking.
> The microphone doesn't give a sine wave.
No, it gives an infinite number of sine waves added up together. Which become a finite number of sine waves after passing through ADC, and then a finite sequence of sine waves after a Fourier transform.
We might not know anything about them in advance, but the patterns are there and could maybe be extracted from the some training data. If only you had a statistical model that was flexible enough to find them…
Validation is then as easy as running the model on some examples outside the training set.
> No, it gives an infinite number of sine waves added up together.
Yeah, and after Doppler shift it is still an infinite number of sine waves - no immediate information gained.
Of course, if there are characteristics in the original noise and its frequency distribution, you could try to find those in the doppler-shifted signal. How would you determine the characteristics? From a dataset of examples, I guess. So now the problem is: recognize a pattern from examples and try to find it in new instances. Sounds like the kind of problem ML has found success in. (If you're now thinking "we don't need ML, just some advanced statistics"… Well ML is often basically a statistical model with lots and lots of parameters.)
Only if you can trust the data gathered from that validation to be representative. You can do that easily when you understand the statistics your model is doing - which is the case with an "old-school" ML solution, but not so with DNNs.
This gets worse the more complex your problem is. I can expect a DNN to pick up the correct frequency patterns in audio time series quickly, as it stands out in the solution space - but with more variables, more criteria, we know it takes ludicrous amounts of data for the model to start returning good results, and it still often fixates on dubious variables.
And then you have to ask yourself - what are your error bars? With a classical approach to estimating train velocity from sound, your results will be reasonably bounded, and won't surprise you. With a DNN, all bets are off.
> How would you determine the characteristics? From a dataset of examples, I guess.
And physics. In this case, a human can apply their understanding of physics to determine what characteristics to expect, verify they exist in the dataset, and encode that knowledge in the solution. A DNN will have to figure this out on its own, and we have no good way to verify it did it correctly (and isn't just overfit on something that's strongly but incidentally correlated).
I agree there are plenty of problems where we don't have a good "first principles" solution - where we're just looking for correlations. DNNs automate this nicely. But such models belong to the category of untrusted ones - they might seem to work now, but because of their opaqueness, we can't treat past performance as a strong indicator of reliability.
> Well ML is often basically a statistical model with lots and lots of parameters.
Yes. But I think it matters if people know what those parameters do.
The main advantage over a physics-based modeling approach - which with enough information, could surely have reached practically 100% accuracy - is that the SVM didn't rely on knowing anything about the location of the access points, or the geometry of the space. The signal strength training data was to be available for free as a biproduct of another device, so this solution had very low cost in the form of manual effort/precise measurement, both of which would have dwarfed a few weeks of intern time.
Given just some microphones picking up whatever sound the train makes on its own, it's not obvious to me that there's a simple solution.
Also, depending on the track used, there may be trains passing by without braking, so you will need at least a classifier to sort these two cases.
I'd argue that using ML to build such a classifier is almost always a time saver.
And if you have the ML pipeline there, why not try to train it to recognize the speed while we are at it? It will likely find out about doppler shift but also do things that would take ages to code manually:
- Use volume levels and volume level differences - Use the clicks at rails junctions to evaluate the speed - Recognize the intensity of the braking/engine running - Use cues like rails vibration at certain speed - Adjust for air pressure difference when it hears the rain
All of that for free. Nowadays, going ML first is becoming a pretty good idea actually.
I guess, such a calculation could have been one of the inputs to the system.
I do get your point that an ML system for such a thing is an overkill. I guess there are more reliable and rugged methods to get the speed of the incoming train (sensors that need not be mounted on the train)
You misunderstood the point of the presentation. The company was a consulting firm that specialized in data science and engineering. Our clients wanted to kick the tires and see what our technical chops were before hiring us but they didn't want to let us use their proprietary and confidential data for our own tech demos.
We didn't want to just use the same open source datasets everyone else did, so we got to thinking about novel datasets we could create that might have applications for industries we sold our services to. From this, the Trainspotting project was born.
Many of us commuted via the Caltrain, which was right next to our office, and we were frequently frustrated with the unreliability (this was in ~2016 or so when car and pedestrian strikes were happening seemingly every week), so we made an app that tried to provide more accurate scheduling.
We used the official API for station:train arrival times, but we found that it was unreliable, so we wanted some ground truth data on whether a train was passing. Since our office was right next to the Castro MTV station, I had the idea to use a microphone (attached to a raspberry pi) to just listen for when the train went by. In addition to ground-truth data for validating arrival times, this gave us a chance to show off some IoT applications. It actually worked pretty well, but it had false positives (e.g. the garbage truck would set it off). So we added a camera.
We pointed it at the tracks and started streaming data off of it. At first we used very simple techniques, processing the raw stream on-device with classic computer vision algos (e.g. Haar cascades) in openCV. We discovered that the VTA, which had a track parallel to the Caltrain and was "behind" the Caltrain in our camera's shot, could cause false positives. Gradually we used more and more complex techniques like deep learning, but the raspberry pi couldn't handle it (IIRC it could only process a single frame in like 6 seconds). So we used a two-stage validation whereby the simpler, faster detectors that could run on the raw stream in real time detected a positive and then we'd send a single frame to run deep learning.
TL,DR: The whole point was to be a tech demo, not to gauge the speed. The trains were either stopping or pulling out of the station, so speed would have been useless.
I also acutely enjoy the notion that a pithy critique of people who refused to simplify the problem they were solving is in itself grossly oversimplified!
I didn't really understand the technology (gradient descent) underlying the training, so I went to grad school and spent 7 years learning gradient descent and other optimization techniques. Didn't get any chances to work in ML after that because... well, ML had a terrible rep in all the structural biology fields and even the best models were at most 70% accurate. Not enough data, not enough training methods, not enough CPU time.
Eventually I landed at Google in Ads and learned about their ML system, Smartass. I had to go back and learn a whole different approach to ML (Smartass is a weird system) and then wait years for Google to discover GPU-based machine learning (they have Vincent Vanhouke to thank- he sat near Jeff Dean and stuffed 8 GPUs into a workstation to prove that he could do training faster than thousands of CPUs in prod) and deep neural networks.
Fast forward a few years, and I'm an expert in ML, and the only suggestion I have is that everybody should read and internalize: https://research.google/pubs/pub43146/ So little of success in ML comes from the sexy algorithms and so much just comes from ensuring a bunch of boring details get properly saved in the right place.
and this is the first part of the description:
“ Machine learning is cool, but it requires data. Theoretically, you can take data from a different problem and then tweak the model for a new product, but this will likely underperform basic heuristics. If you think that machine learning will give you a 100% boost, then a heuristic will get you 50% of the way there.
(..)”
https://developers.google.com/machine-learning/guides/rules-...
Basically the heading becomes the random discussion topic that gets thrown in the room.
Maybe there is an experimental social platform in that:
(Re-)create a HN or reddit look-alike, but instead of user submitted links just pick random headings from news sites. Every ten minutes, post a new one without any context or link to be discussed and voted by the audience.
No idea where this would take us.
So you are basically asking for a subset of HN? To avoid echo chamber I think the links is a good thing.
The problem is that the kind of ML that involves downloading a framework from github and tweaking features until the percent goes up is actually built on certain statistical models under the hood that people don't understand and that don't fit the process they're trying to model.
When the statistical model is correct you don't need loads of data. E.g., you don't need more than a thousand respondents to make valid inferences about millions of people in a sociological survey.
People are downloading ready-made models from repositories to try and solve minor problems.
Guess what, your problem might be a simple linear regression. Yes you can solve it with a DNN (one level, one neuron - but hey, don't keep that from putting it into your CV) but you don't need to.
It was really cool. Attempting to implent Hamiltonian MCMC on a single neuron really forced you to learn what a gradient is in regards to NN.
Using a pre-trained model only works if the usecase it was trained for matches you're closely enough.
I fine tuned YOLOv5 with a few dozens hand-labelled images to make an object detector in a semi-controlled environment.
The idea that you need a million images to train a detector or a classifier is now totally wrong. Fine-tuning can be done on a very small dataset.
From the scikit-learn faqs:
> Will you add GPU support?
> No, or at least not in the near future. The main reason is that GPU support will introduce many software dependencies and introduce platform specific issues. scikit-learn is designed to be easy to install on a wide variety of platforms. Outside of neural networks, GPUs don’t play a large role in machine learning today, and much larger gains in speed can often be achieved by a careful choice of algorithms.
Of course, there are libraries that can support GPU acceleration for numpy calculations using matrix transformations now. Nonetheless, they are not often necessary.
Early in my career I moved to Silicon Valley to work for a large company. The project was a machine learning project. I was taking models defined in XML, grabbing data from a few different databases, and running it through a machine learning engine written in-house.
After a year and a half, it came out that our machine-learning-based system couldn't beat the current system that used normal statistics.
What rubbed me the wrong way was that the managers brought someone else in to run the data, manually, through the machine learning algorithm. More specifically, what bothered me was that we didn't attempt this kind of experiment early in the project. It felt like I was hired to work on a "solution in search of a problem."
Career lesson: Ask a lot of questions early in a project's life. If you're working on something that uses machine learning, ask what system it's replacing, and make sure that someone (or you) runs it manually before spending the time to automate.
For example, have you ever tried to autodetect a published datetime of a news article published online? In many cases, it will be in metadata, or in the time/datetime tag.
However, there still many websites where published time is just written somewhere with no logic at all.
Writing a RegEx script by hand can resolve a problem. But every time I speak about it with our clients/prospects, they ask about ML that we use to parse news content.
Product: https://newscatcherapi.com/news-api
The question is then, (1) do we have datasets that are highly representative of the solutions to our problems? and (2) are our current systems sensitive to the relevant variations in these datasets?
If (1) is NO, then ML is impossible. If (2) is YES, then it's unlikely to provide a big ROI.
For example, processing images according to some trained model instead of fixed rules and formulas introduces the risk of mismatched models (e.g. landscape photographs treated as line art from anime). Cases like self-driving cars not seeing obstacles are more obvious and more tragic.
I like to think of them as forgiving sieves of patterns in data.
Overfitting a sieve will exclude a large number of almost positive cases, loose fitting will include a large number of mostly negative cases.
And there is always a danger of falling into a local minima and not being able to come out of it.
In the community, there is a trend that "complicated == better". imho, more is less in industrial ML. You need to deal with model management, worry about inference & latency when the model gets bigger. The author has another article where he argues that data scientists need to be full stack ninja. While I don't fully agree with that statement, I think it benefits the company in many many ways. Data scientists need to meet engineers in the middle, and all these challenges need to be considered from day 1. Another trend I see is that some data scientists are not driven by the question "Can we solve this problem for the company?", but rather "Can we solve this problem using ML/DL?". This will lead data scientists to use the shiny and trendy models, even if it is not suitable for the job. I would blame management here, in some environments, data scientists are evaluated based on "fancy" models they build, not solutions that they provide. Solutions can be simple (but not simpler) rules.
Rich Sutton's "bitter lesson" says the weight will move in time in favor of ML. http://www.incompleteideas.net/IncIdeas/BitterLesson.html
This is why ML is pretty much the only option for autonomous driving, but not for calculating credit scores.
I know that was probably an offhand example, but it's illustrative of the kinds of non-functional requirements that can make ML solutions more or less viable as soon as the technology has contact with human society.
First rule of optimization: Don't optimize first. First rule of automation: Don't automate first.
The first rule of everything should be “it depends”.
Reminds me of the book "Everything is obvious", where they experimented a few times and showed that in complex systems, advanced prediction systems made on many available and seamingly relevant variables are only marginally better (2 to 4% in the experiments) than the simplest heuristics you can use. They interpreted that as a limit of predictability, because systems with sufficient complexity behave with a seemingly irreducible random part.
The Second Rule of Machine Learning - Start Machine Learning with simple shallow models.
50% of the problems are solved with good data choice of data + some generalized linear model.
30% remaining problems solved with shallow models or old school ML models. Anything from support-vector machines, decision trees, nearest neighbors, very shallow neural networks.
Remaining 20% require more work.
I myself tried several time to “not use ML” for some easy computer vision tasks where traditional CV methods are supposed to work. Well I always end up in situations where they don’t work well without fine parameter tuning, and tuning the parameter for a situation breaks the model in other situations, so you start adding layers of complexity to automatically tune the parameters, but the parameter tuning system also has its own parameters... While a simple neural net is trained easily and is much more robust, saving a lot of time and complexity.
Another proof of that is that CV products only started meaningfully entering the market after ML became applicable to CV (after 2015 for complex tasks, or earlier for simpler stuff like MNIST).
That being said, you are right to say that ML changed the game entirely for CV in industry at large.
It revolutionized the field in a little over a decade and brought forth new frontiers.
I do believe that CV is an area ML will excel well into the next century.
Perhaps, we will find a way to chain together ML systems dynamically, overseen by a procedural system that makes real time decisions in understanding its input.
Once you have deep understanding of your domain and system, finding the places where ML truly adds value is a lot easier. Also, you'll have a basic understanding of how things are without it and you'll know whether it is working better or not and whether that's worth the trouble.
My project is powered by AI.
Back then, when I did social research at university, I found it helpful to just look at the raw data. This is immensely helpful for familiarizing yourself with the data and discerning patterns that high-level analysis wont reveal easily. (In this case, you may want to start with a subset for evident reasons.)
Once you think you have a market, you should see how it can be done FIRST without your "fancy miracle technology" because NOBODY buys a technology because it's sexy or trendy: they buy because it added value in terms of more capability or lower costs.
And ALL problems have current solutions that almost certainly DO NOT use anything as complex as your technology solution so you have to trend very carefully and deliberately in a rational sense: what value are we REALLY adding? That starts with knowing your competition and the current solution to solving the problem first and then finding every reason why your technology won't work or will be problematic.
You ONLY have market potential once you've exhausted those faults or have objective arguments for your value proposition that have been validated by actual customers. The actual prove is made when they are willing to write a PO to you for the solution. Until then, everything you are doing is unproven.
* 1 -> the discovery of specific families of non-linear classification algorithms (with image and language patterns being examples succesful new domains). the domain where these approaches are productive might be significantly smaller than what all the hyperventilation and obfuscation suggests.
* 2 -> the ability to deploy algorithms "at scale". this cannot be overemphasized. Statistics used to be dark art practiced by scienty types in white lab coats locked in ivory towers. With open source libraries, linux, etc to a large degree ML means "statistics as understood and practiced by recently graduated computer scientists"
* 3 -> business models and regulatory environments that enabled the collection of massive amounts of personal data and the application of algorithms in "live" human contexts without much regard for consent, implications, risks etc. Compare that wild west with the hoops that medical, insurance or banking algorithms are supposed to pass
Conclusion, ML is here to stay in some shape or form, but ML hype has an expiration date
> First, design and implement metrics.
> Before formalizing what your machine learning system will do, track as much as possible in your current system. Do this for the following reasons:
> * It is easier to gain permission from the system’s users earlier on.
> * If you think that something might be a concern in the future, it is better to get historical data now.
:-/
Even with a clean dataset, most clients will want basic arithmetic calculations: averages, counts, percentages, standard deviation, etc. Occasionally they'll want some basic logistic models, something slightly more causal. If they go straight to machine learning without these steps, do they actually understand their problem and what they want? Or are they reaching for the shiniest thing they've heard of?
Unfortunately most decision making C-suites are not engineers who fall for the marketing hype and burn through time and capital without tangible outcomes.
It was a few years ago. I had to classify pictures of closed and opened hands. I thought surely I don't need ML for simple stuff like that: a hue filter, a blob detector, a perimeter/area ratio should give me a first prototype faster and given the little amount of data I had (about a hundred images of each), not worth the headache. I quickly had a simple detector with 80% success rate.
Then as I was learning a new ML framework, I tried it too, thinking that would surely be overengineering for a poor result. I took the VGG16 cat-or-dog sample, replaced the training set with my poorly scaled, non-normalized one, ran training for a few hours and, yes, outperformed the simple detector that took me much longer to write.
Now in computer vision, I think it makes sense to try ML first, and if you are doing common tasks like classification or localization of objects, setting up a prototype with pre-trained models has become ridiculously easy. Try that first, and then try to outperform that simple baseline. In most case, it will be hard and instead worth improving the ML way.
To clarify, when I talk about ML I'm primarily referring to classifier algorithms and approaches (including nlp). In the large part the ML is being used to generate classifier rules which generalize patterns, and data cubes are often used to look for aggregations and data sequences which generalize patterns. The problem is that random patterns happen all the time, and may even persist for a long time despite a lack of real correlation. Semantic analysis of data cube output is really important in order to find meaningful patterns.
What I'm getting at is I often wonder why most ML projects try to treat it like it's magic. Human assisted learning has shown repeatedly to be the system which actually works in practical application. The classifier output needs to be pruned to remove rules that only held true in the sample data, or were merely coincidental, or simply have no practical value.
Approaches like this are not cheap to set up and may in the end still only produce the same results as the existing entirely non-ML based system. What is the likely scale of work compared to the benefit is the first question I ask myself before working on anything. If I don't have objective data to answer that you have to do some research to find out. Never try to build a massive or complicated system you don't have objective reasons to expect will be worth the effort. That's precisely what people have been doing with ML constantly. It's little wonder most developers have such low opinions of ML projects.
It's true that if your rules grow in complexity, this might make it harder to maintain, but the good thing about rules is that they tend to be fully explainable, and they can be encoded by domain experts. So the maintenance of such a system does not need to be done exclusively by an ML engineer anymore.
Here is where I insert my plug: I have developed a tool to create rules to solve NLP problems: https://github.com/dataqa/dataqa
In the past I attended several meetings with customers where I was actively discouraged asking questions which would help us deliver a good meaningful solution as long as the customer would be happy "investing in a ML solution". And they were...
They just want to do ML and are looking for a problem that can be solved with it. Then they will likely ignore you when you say this problem has also neat traditional solution.
This is further exacerbated by corporate actions like competitions for best AI (or Blockchain, etc.) project. Which you typically can't participate in if you have traditional solution even if it is way better.
Example: https://beta.openai.com/examples/default-translate
They even use flawed results in their marketing materials that they didn't validated with domain experts. ("Où est les toilettes ?" is not french).
If on the other hand you start from having some statistics to do on the data you have, then you might at some point find yourself doing the sexy subset of it that we call 'ML', and fine.
A few thoughts on how to maximize your chances of winning in this case:
https://medium.com/criteo-engineering/making-your-company-ml...
Your input value is X, you multiply it by your slope and add your intercept to get the output (the Y value on the line).
The 'training' of an ML algo is really just finding the line-of-best-fit so that you can make predictions. So your line-of-best-fit is encoded in these two parameters, allowing you to make predictions about what the output would be for arbitrary input.
The problems people are throwing at ML have many more parameters and dimensions, but the training is a matter of finding those parameters that come closest to predicting the outcome. The 'model' is this set of parameters that allows the function to make predictions.
(disclaimer: also an ML noob, correct me if I'm wrong)