Dear AI startups: Your ML models are dying quietly
sanau.co
sanau.co
No system should ever fail silently if a required field suddenly goes missing or has the wrong type/unit.
E.g., give rows a type, give columns a type, and then if you multiply two matrices, see if the type of the columns of the first matrix and the type of the rows of the second matrix match.
Also, the typesystem could use units, just as they are used in physics.
http://www-ksl.stanford.edu/knowledge-sharing/lib/unit-conve...
Given the way this appears to be defining some sort of class-based hierarchy, with the correct type requirements this code should perform exactly the sort of unit/type checking the grandparent post wanted, but I didn't go digging into the rest of the project files to confirm this code defined things in the correct manner for that to work.
You could even do it with SQL field (column) names and natural joins. The performance hit is even bigger and discipline/retraining is even bigger, as few SQL users are used to thinking about column names (rather than types) as type definitions.
Further, this would only catch a class of bugs that are relatively easy to spot manually. The difficult case is when the processes that generated the numbers have slowly changed in time (e.g., you have a very different mix of visitors/customers now than during model training).
We realize that it's really important for data science, product management and engineering teams to discuss and ideally monitor any new or any changes in data capture and processing that feeds into a machine learning model.
https://www.virustotal.com/#/url/e4accf1e046c8266168b9038763...
1. The model uses click-through data as an input. Your frontend engineer moves the UI element being clicked upon to a different portion of the page for a certain category of results. This changes the baseline click-through rate. The model assumed this feature had a constant baseline across all results, so the new feature value now needs to be scaled to account for the different user behavior. Nobody thinks to do this.
2. The frontend engineer removes a seemingly-wasted HTTP fetch to reduce latency. This fetch was actually being used to calibrate latency across different datacenters, and was a crucial input to a data pipeline to a system of servers (feeding the ML model) that the frontend team didn't control and wasn't aware of.
3. The frontend engineer accidentally triggers a browser bug in IE7 (gimme a break, it was 9 years ago) that prevents clicks from registering when RTL text is mixed with LTR. Click-through rates decline precipitously in Arabic-speaking nations. This is interpreted by an ML model as all results being poorly performing in Arabic countries, so it promptly starts cycling through results, killing ones that had shown up before with no clicks.
4. A fiber cable is cut across the Pacific. This results in high latency for all Chinese users, which makes them abandon their sessions. A ML model interprets this as Chinese people being less interested in the news headlines of that day.
5. A ML model for detecting abusive traffic uses spikes in the volume of searches for any one single query over short periods of time as a signal. Michael Jackson dies. The model flags everyone searching for him as a bot.
6. A ML model for search suggestions uses follow-up queries as a signal. The NYTimes crossword puzzle comes out. Everybody goes down the list of clues and Googles them. Suddenly, [houston baseball player] suggests [bird sound] as related.
It's just well-known things being rediscovered because people are treating this as a new field - but everything mentioned in this article would be the same for company tracking some other sales efficiency metric twenty years ago, except that the "change in front end framework" would be "change in preorder sales reporting templates", causing very similar problems in your data analysis.
I've used these solutions, but while Jupyter is amazing (my startup https://kyso.io started as a way to share notebooks) I'm not sure if deducing which cells to convert into an endpoint is the way to go - especially since you will need to also host model files and extra data?
That's why production ML systems should monitor for data drift, model accuracy, and a host of other factors that may not be obvious at first. That's part of what TensorFlow Extended does.
See also: "What’s your ML Test Score? A rubric for ML production systems" https://ai.google/research/pubs/pub45742
If the model's ROC-AUC falls by 0.1, the mean/s.d. of one of it's inputs changes by more than 50% or the number of N.a.s for an input increases suddenly or if the monitoring report dies, then the model owner should get an alert quickly.