376 karma · joined June 9, 2018
My comment was too judgemental. People should be allowed to say that they enjoy something without any follow-up and without being judged for it. I think that sometimes it just seems like a facade when people say they really like math because when you try start a conversation on the topic its like they're not actually interested in it at all. It gives the impression there's something disingenuous about their proclamation of liking math. But perhaps its just the way I personally have approached it.
Its hard to read past nonsense like this.
> it's really hard for me, as a non-expert, to assess which papers are true advancements
Its hard for me too, though I wouldn't consider myself an expert, just someone with a moderate amount of experience. Learning to discriminate important from less-important papers is another skill which takes effort to develop.
The whole objective here is personal learning and this advice would be wildly different for how to practice ML professionally. The approach is directly analogous to advising a beginner programmer to get better at programming by actually writing computer programs.
> Most of the cutting edge papers are trained on several $100k worth of GPU time
Its besides the point, but I said nothing about a requirement that the methods that you choose to implement and learn from having to be cutting edge. More to the point, unless we have a different definition for what "cutting edge" means, you're wrong that "most of the cutting edge papers" require high computational resources. If that were true it would be nearly impossible for the field to make progress at the pace it does. There is a plethora of research in purely algorithmic approaches which do not require massive compute resources, and in fact this is the most productive portion of research to learn from because there the focus is on theory and progress in how to conceptualize / frame ML problems. Works which amount to "we took method X and massively scaled it up" are (in my opinion) less intellectually interesting to someone seeking to grow their knowledge in ML (though the results may might be extremely impressive and impactful, and it may be intellectually very interesting for the working directly on that project).
> How can you be sure that your implementation is correct, if you can't train it (hence you can't run proper inference with a good model)?
This is like asking how you can be sure that you've correctly implemented a B tree if you haven't used it to serve a distributed database to 1 million users. The answer is small isolated tests.
One of the best ways to really test your knowledge of an ML algorithm is to design and write unit tests to assert it behaves correctly on trivial cases. You'll find bugs in your implementation, but you'll also be forced to think carefully about what the core characteristics of the algorithm are that must be asserted in order to convince yourself that its correct. Its a common beginner mistake in ML to just run/train your model and have that be the only test of its correctness. Its like deploying a web service with zero tests and letting "do I get X number of users" be the only test of your code's correctness. It sounds insane but its basically equivalent to what most beginners do in ML (my former self included).
- Score-Based Generative Modeling through Stochastic Differential Equations https://arxiv.org/abs/2011.13456
- Structured Denoising Diffusion Models in Discrete State-Spaces https://arxiv.org/abs/2107.03006
- Efficient and Modular Implicit Differentiation https://arxiv.org/abs/2105.15183
- Scalable Gradients for Stochastic Differential Equations https://arxiv.org/abs/2001.01328
- Bayesian Optimization with Unknown Constraints https://arxiv.org/pdf/1403.5607.pdf
- SE(3)-Transformers: 3D Roto-Translation Equivariant Attention Networks https://arxiv.org/abs/2006.10503
- DiffDock: Diffusion Steps, Twists, and Turns for Molecular Docking https://arxiv.org/abs/2210.01776
You can view this approach in the same way that a beginner learns to program. The best way to learn is by attempting to implement (as much on your own as possible) something that solves a problem you're interested in. This has been my approach from the start (for both programming and ML), and is also what I would recommend for a beginner. I've found that continuing this practice, even while working on AI systems professionally, has been critical to maintaining a robust understanding of the evolving field of ML.
The key is finding a good method/paper that meets all of the following
0) is inherently very interesting to you
1) you don't already have a robust understanding of the method
2) isn't so far above your head that you can't begin to grasp it
3) doesn't require access to datasets/compute resources you don't have
of course, finding such a method isn't always easy and often takes some searching.
I want to contrast this with other types of approaches to learning AI with include
- downloading and running other people's ML code (in a jupyter notebook or otherwise)
- watching lecture series / talks giving overviews of AI methods
- reading (without putting into action) the latest ML papers
all of which I have found to be significantly less impactful on my learning.
The notion of a "rut", by contrast, imposes a physical structure on the concept (think is a person who is literally stuck in a ditch in the ground) which doesn't lend itself to your thinking of it as "inevitable" for you to ever get out of it.