That’s such a fascinating background to have! It must be strange to retire and, far from the cliché of your skills having been made obsolete and irrelevant, instead you’re an expert on what’s probably the forefront of modern technology.
If you can spare a moment to answer: is there any knowledge from the 80s neural net ‘summer’ which you think has been forgotten now? As someone who’s concerned by both poor performance and overfitting (strongly correlated), I thought it was a shame that lots of the research around pruning (optimal brain damage, optimal brain surgeon, etc) has been forgotten[0].
[0] It feels to me as though the ML community has convinced itself - in defiance of information theory - that highly overparameterised models are totally OK, and can even successfully extrapolate with greater than random accuracy in the general case (‘double descent’). I worry that lots of these models are effectively succeeding only at interpolation problems, and only by virtue of the massive hardware advances that let us memorise the entire training set in these enormous models - basically Runge’s phenomenon writ large. People seem to be convincing themselves of magical things that are not mathematically sound.