Grokfast: Accelerated Grokking by Amplifying Slow Gradients
arxiv.org
arxiv.org
I remember seeing grokking demonstrated for MNIST (are there any other non synthetic datasets for which it has been shown?), but the authors of that paper had to make the training data smaller and got a test error far below state of the art.
I'm very interested in this research, just curious about how practically relevant it is (yet).
Improving generalization in deep learning is a big deal. The phenomenon is academically interesting either way, but e.g. making sota nets more training data economical seems like a practical result that might be entirely within reach.
however, techniques we learn from grokking about implicit regularization might be helpful for the training regimes we actually use
I'm not so sure. Reasoning is the next big hurdle, and grokking and parametric memory seem very effective here.
all that grokking really means is that the 'correct', generalizable solution is often simpler than the overfit 'memorize all the datapoints' solution, so if you apply some sort of regularization to a model that you overfit, the regularization will make the memorized solution unstable and you will eventually tunnel over to the 'correct' solution
actual DNNs nowadays are usually not obviously overfit because they are trained on only one epoch
1) Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs https://arxiv.org/abs/1802.10026
2) Visualizing the Loss Landscape of Neural Nets https://arxiv.org/abs/1712.09913
There's also a very interesting body of work on merging trained models, such as by interpolating between points in weight space, which relates to the concept of "basins" of similar solutions. Skim the intro of this if you're interested in learning more: https://arxiv.org/abs/2211.08403
Reviewing the literature, I see the concept is more commonly referred to as "flat/wide minima"; e.g., https://www.pnas.org/doi/10.1073/pnas.1908636117
Both use averaging.
You have countless LLMs, use one of them to generate new names that don't require ten billion new disambiguation pages in Wikipedia.
MNIST is solvable using two pixels. It shouldn’t be one of two benchmarks in a paper, again just in my opinion. It’s useful for debugging only.
really? do you have any details?
agree it has no business being in a modern paper
Training a GPT-2 sized model costs ~$20 nowadays in respect to compute: https://github.com/karpathy/llm.c/discussions/481
https://arxiv.org/abs/2405.19454
It's a bit surprising that apparently this hasn't been done before - to see how universal of a phenomenon this is.
Even figuring out how to induce grokking behavior on a 100M model or OpenWebText would be a big leap in the understanding of grokking. It's perfectly reasonable for a paper like this to show results on the standard tasks for which grokking has already been characterized.
Academic research is often times about making incremental steps and limiting uncertainty.
Making something work for MNIST is already so much work, researchers don’t have the time, money, and energy to run experiments for 10 datasets.
Complex datasets are much harder to get a proper model trained on due to increased complexity - larger images, tasks, classes, etc.
Also, as soon as you run your experiments on more datasets, you create an opportunity for reviewers to take you down - “why didn’t you test it on this other dataset?”