Pen and paper exercises in machine learning (2021)
arxiv.org
arxiv.org
>We may have all heard the saying “use it or lose it”. We experience it when we feel rusty in a foreign language or sports that we have not practised in a while. Practice is important to maintain skills but it is also key when learning new ones. This is a reason why many textbooks and courses feature exercises. However, the solutions to the exercises feel often overly brief, or are sometimes not available at all. Rather than an opportunity to practice the new skills, the exercises then become a source of frustration and are ignored.
Typical exercise soltion: How to draw an owl. 1. Draw some circles. 2. Draw the rest of the $@#% owl.
Another sad thing is that we lose our memory in a step function. We remember something for a long time without using it, and then all of sudden it's gone from our memory. That may explain why many lifers in Google could interviewed so badly. They joined Google when there was no leetcode and when Google had amazing high standards in asking math and algorithm puzzles and novel systems design questions. So, they are really good. They are also confident because they achieved a lot and led amazing projects in Google. Yet when they started to answer interview questions, they struggled with basic facts.
Sort of makes one wonder how important those “basic facts” really are.
That's not how memory works, in fact the more we recall something the more we start misremembering associated details
Not sure memory loss is quite that binary. Heading for mid-40s here, and often things are still there, but it's higher latency to fully recall it. A bit like it's stored off in Amazon Glacier. Or often I can get the first byte quickly ("I know that person's surname starts with a P") but retrieving the full answer takes longer and some brute force iterations. It's as if I've hit a hash-table collision and need to binary-compare the results, or I've reached a node in a tree-structure that has many branches ("Is it Pfeffle? Piper? No, Pfeiffer!")
How was that any better? Leetcode is math and algorithm puzzles.
That is - drawing things on paper, no formulae. I used to do similar exercises with the printed Iris dataset and giving people a transparent foil to draw the classifier. Then, giving them another sheet of paper with the validation dataset. The exercise was loved by people from high-school students to managers.
I developed an interactive version, https://github.com/stared/which-ml-are-you. Unfinished as a game, but it works.
Digital pdf is a great way to break free of these limits and this document is impressive indeed with the detailed solutions.
- Everyone “knows” that the Adam optimizer’s proof is incorrect, but we still use Adam because we don’t want to redo hyperparameter search with a different optimizer that’s proven to converge but probably performs worse.
- Everyone “knows” that the Wasserstein loss for GANs has a better convergence proof, but nobody uses it because the generated images look like crap compared to what you get from stylegan* with their default config.
It’d be nice if ML proofs led to better performance, but that’s not often the case. I see far more progress from better data preprocessing and from bringing in knowledge from other fields like signal processing.
as an ee phd can you derive basic control theory or electromagnetic relationships?
Only comment is that it seems quite heavily focused on graphical models, more bayesian/nn concepts would be great to see in this!
I know there is a more elegant approach that makes better use of the determinant's core properties, but it's been ages since I saw it.
https://math.stackexchange.com/questions/38701/how-to-calcul...
Though, there are other ways to derive it. My personal opinion is that the vector and matrix calculus derivations in the book are too verbose, but this style may be more comfortable for some readers. My personal opinion is that the semidefinite and cone optimization communities have more concise ways of deriving these kind of derivatives and relationships. For example, this can be seen in Boyd and Vandenberghe's Convex Optimization or Ben-Tal and Nemirovski's Lectures on Modern Convex Optimization.
Brute-forcing it would be to write down the multivariate polynomial in full generality, as there would be no way to break it up using the logarithm.
For example, https://arxiv.org/pdf/1909.03562.pdf gives version 5 of https://arxiv.org/abs/1909.03562, as noted in the left margin of the first page.
Reading the first few sections, it seems that the ideas are there - especially in the proofs - plenty of motivating ideas, and the kind of "raw index crunching" that the paper begins with gives way to more ideas. Doubters might read section 1.6 about the power method for finding the largest eigenvalue. It convinced me that the ideas were worth reading.
Yeah 1.6 is a really cool exercise! https://arxiv.org/pdf/2206.13446.pdf#page=19
It's so cool to see of why this works (as an engineer I learned about power method with handwaving explanation "it works in the limit" but I never knew why it works).
So what do we do if we want u_2, the eigenvector that corresponds to lambda_2 ? Math overflow says we can just subtract the u_1 subspace from A [1] then repeat, but would that be numerically stable? (i.e. will that work with floats?)
[1] https://math.stackexchange.com/questions/1114777/approximate...
If you don't understand the purpose of proofs, then this resource is not aimed at you.