I hoped that it would be a little open ended as most questions in ML in real life are open ended.
What would take this repo to the next level is to have a reproducible data generation function for each exercise as well as a reasonable metric which must be passed. I don’t see anything that requires my classification auc to be over 0.5 which would be a basic criteria of bug-free code.
I was reverse engineering the ML interview pipeline for myself and that's how I stumbled upon all this.
I think the data aspect does make sense tho. I might add that as the next thing to do