One of the key insights I took away was the importance of using out-of-sample predictive accuracy as a metric for regression tasks in statistics—just like in ML. The standard best practices in STATS 101 is to compute R^2 coefficient (based on data of the sample), which is akin to reporting error estimates on your training data (in-sample predictive accuracy).
IMHO, statistics is one of the most fascinating and useful fields of study with countless applications. If only we could easily tell apart what is "legacy code" vs. what is fundamental... See this recent article https://www.gwern.net/Everything the points out the limitations of Null Statistical Hypothesis Testing (NHST), another one of the pillars of STATS 101.