As for the approaches - it seems the papers you mention do this statically, and din't produce an algorithm for actually speeding up computations with their 20-50% results - which was a large part of the difficulty. I'll probably take some time off one day and do a thorough review of the literature finally.
Ultimately I want to add a citations page with these and some other papers people have posted in the comments soon. I expect sooner or later someone will find this algorithm written up by someone :)
While development asked gpt-4 and googled trying to find information about this method - everything seemed either static, or focused on removing full dimensions/layers ad hoc with eventual retraining. Didn't find anything matching this idea exactly.