One of the ML mistakes that will hurt most: not collecting the right features
twitter.com
twitter.com
Anyone can throw a mess of features into a huge hog of a model and get some sufficiently accurate input. But with feature engineering, you essentially absolve the computational model from needing the vast majority of the capacity in the naive case, instead "pre-computing" those transforms which would otherwise be learned by the model. With the reduced need for capacity, the ability to refine the model for interpretability's sake (and optimization stability) also improves.
Of course part of success is also collecting the right data, but focusing on the right features is a slightly more nuanced and focused way to look at it IMO.
I spent a month writing and implementing a new company playbook on a building products with data science.
It boils down to including data science early during discovery (In part to help identify what data we have available up front), and viewing them as a squad stakeholder rather than a team for one-off work. It also pushed everyone to start thinking about the reliability of data for every product we release.
Some of the databases I’ve seen at companies is just an absolute abomination.
I've seen examples of where good and careful data collection in advance would have solved the problem, but instead, someone decided to transform the (easiest) collected data into a format suitable for DL - and spent months trying to optimize and tweak the DL model. Because their rationale was that it's much easier to just fine tune an out-of-box DL / transfer learning model, than spending ages on extra planning, data collection, feature engineering, and what not.
The word 'information' is doing some heavy lifting here. There is nuance - what information could be learned or computed from the data? We can only really guess. When we're wrong about this, then we will design models imperfectly.
Is there enough information in internet text scrapes to solve unseen high school mathematics word problems? The feature engineers of olde would have said no. But there is, to a surprising degree.
(If it's wrong, the scope of the mistake reaches far beyond this article.)
Why does the second set of features work while the first doesn't? I like to think of the second set having a model of causality, to figure that out you need some domain expertise and experience. Without that expertise, the feature set of non-causal factors is almost infinite.
The classic example is the "air conditioned room" thought experiment. Imagine feeding an AI model a dataset with a bunch of datapoints like this
Room hot | AC on full blast
Room cold | AC off
Room warm | AC on 50%
Then asking the AI how to make a hot room cold. The AI is very likely to say "turn off the AC" because "AC off" is a feature of cold rooms.
Maybe we think that if we give the AI more information, it could do better. Lets imagine that our room has a window and an oversized AC. We'll now be feeding it info about the source of heat as well as the source of cold, so surely it will get the right answer this time:
Window open | AC on full blast | Room cold
Window closed | AC off | Room cold
Window 50% | AC on 50% | Room cold
This dataset is going to have the AI thinking that neither the window nor the AC impact room temperature, and that you might be able to close a window by turning off the AC.
The key thing to understand is that control loops change the "physics" of your system and that knowledge about the control-physics does not transfer to the uncontrolled-physics.
The fact there is a control loop here doesn't actually seem important if you represent the features correctly.
I'd argue that data representation is actually what you need to get right (ie, in this case the behavior in the time domain).
What machine learning algorithms even try to determine causation?
1. data
2. algo
https://twitter.com/melling/status/1537562264387166209?s=21&...
People make the same mistake in their decision process.
I was going to add that his model wasn’t explainable either but only so much geek humor…