HNHacker News
TopNewBestAskShowJobs

charleshmartin

83 karma · joined July 5, 2015

Founder, Calculation Consulting Machine Learning and Data Science Expertise

http://www.linkedin.com/in/charlesmartin14

http://calculationconsulting.com

submissionscomments
charleshmartin··on Detecting Overfit Layers without any Data
If you train a model for too long, it may overfit it's training data. Not surprising, this has been know for like forever. But did you know you can detect the signatures of overfitting in the layer weight matrices directly, without needing access to any data (train or test) ?

In our recent paper (with hari kishan prakash ), - : - , we show this explicitly in 2 different classic grokking experiments. And the overfitting we see is very different from what has been seen before!

paper: https://arxiv.org/abs/2602.02859

charleshmartin··on Attention at Constant Cost per Token via Symmetry-Aware Taylor Approximation
Right. If the dynamics of training are governed by RG flow, then the best optimization path should remove redundant directions, as specified by the RG operator(s)
charleshmartin··on Setol: SemiEmpirical Theory of (Deep) Learning
We present a SemiEmpirical Theory of Learning (SETOL) that explains the remarkable performance of State-Of-The-Art (SOTA) Neural Networks (NNs). We provide a formal explanation of the origin of the fundamental quantities in the phenomenological theory of Heavy-Tailed Self-Regularization (HTSR): the heavy-tailed power-law layer quality metrics, alpha and alpha-hat. In prior work, these metrics have been shown to predict trends in the test accuracies of pretrained SOTA NN models, importantly, without needing access to either testing or training data. Our SETOL uses techniques from statistical mechanics as well as advanced methods from random matrix theory and quantum chemistry. The derivation suggests new mathematical preconditions for ideal learning, including a new metric, ERG, which is equivalent to applying a single step of the Wilson Exact Renormalization Group. We test the assumptions and predictions of SETOL on a simple 3-layer multilayer perceptron (MLP), demonstrating excellent agreement with the key theoretical assumptions. For SOTA NN models, we show how to estimate the individual layer qualities of a trained NN by simply computing the empirical spectral density (ESD) of the layer weight matrices and plugging this ESD into our SETOL formulas. Notably, we examine the performance of the HTSR alpha and the SETOL ERG layer quality metrics, and find that they align remarkably well, both on our MLP and on SOTA NNs.
charleshmartin··on Setol: A SemiEmpirical of (Deep) Learning
SETOL is a theory of NN layer convergence. It argues that the individual layers of NN converge at different rates, and the 'Ideal' state of convergence can be detected simply by looking at the spectral properties of the layer weight matrices.

SETOL derives the weightwatcher HTSR layer quality metrics (alpha, alpha-hat) from first principles using techniques from statistical mechanics and quantum chemistry.

SETOL also shows that when an NN layer is 'Ideal', it satisfies the so-called Wilson Exact Renormalization Group condition (called the TraceLog condition)

In other words, SETOL provides empirical layer quality metrics that can be used to determine how well a model is trained or fine-tuned. It can help AI engineers get their AI models to the best state possible.

And while this work uses techniques from theoretical physics and chemistry, you don't need know any mathematical physics to read the paper; it is fully self-contained.

Most importantly, all of the experiments are 100% reproducible, and you can test the theory yourself on your own AI models using the open-source weightwatcher tool.

charleshmartin··on DyLoRA: Parameter Efficient Tuning of Pre-Trained Models
There are good theoretical reasons behind this as well https://calculatedcontent.com/2023/02/01/deep-learning-and-e...
charleshmartin··on Ask HN: Why so much hatred towards OpenAI?
Envy
charleshmartin··on Ask HN: Why so much hatred towards OpenAI?
ENVY
charleshmartin··on Weightwatcher: Data-Free Diagnostics for Deep Learning
WeightWatcher (w|w) is an open-source, diagnostic tool for analyzing Deep Neural Networks (DNN), without needing access to training or even test data. It is based on theoretical research into Why Deep Learning Works and uses the Theory of Heavy-Tailed Self-Regularization (HT-SR), published in JMLR and Nature.
charleshmartin··on Neural Network Loss Landscapes: What do we know? (2021)
https://calculatedcontent.com/2015/03/25/why-does-deep-learn...
charleshmartin··on The Principles of Deep Learning Theory
What can they predict with the theory ?
charleshmartin··on A new link to an old model could crack the mystery of deep learning
Oh and here's a 2 hour deep dive into the theory https://bluejeans.com/playback/s/cyQnC9EZZSB9HlHDjB1Zpu3fs0J...
charleshmartin··on A new link to an old model could crack the mystery of deep learning
agreed!
charleshmartin··on A new link to an old model could crack the mystery of deep learning
That's fine. Here's a few of my online talks:

UC Berkeley / ICSI: https://www.youtube.com/watch?v=6Zgul4oygMc

Stanford ICME: https://www.youtube.com/watch?v=PQUItQi-B-I

and a couple of Mikes's

Institute for Pure & Applied Mathematics (IPAM) : https://www.youtube.com/watch?v=fmVuNRKsQa8

Physics Informed Machine Learning: https://www.youtube.com/watch?v=eXhwLtjtUsI

and our KDD workshop KDD: https://dl.acm.org/doi/abs/10.1145/3292500.3332294

many more talks and papers available

And my favorite, a podcast that has featured LeCun himself:

https://blog.rebellionresearch.com/blog/theoretical-physicis...

charleshmartin··on A new link to an old model could crack the mystery of deep learning
Here's an alternative approach, that actually provides real world results

https://calculatedcontent.com/2019/12/03/towards-a-new-theor...

Using techniques from statistical mechanics and strongly correlated systems, we can compute the average-case-behavior of a real world DNN

We believe we can reproduce some of the results of the NTK by using a Gaussian Random Matrix. But if we use a more realistic, heavy tailed matrix, we get more practical results

A early fork of the theory has been published in JMLR https://arxiv.org/abs/1810.01075

and the empirical results in Nature Communications https://www.nature.com/articles/s41467-021-24025-8

and we have an open source tool , weightwatcher, which can be used in production

pip install weightwatcher

https://github.com/CalculatedContent/WeightWatcher

Please give it a try and let me know if it is useful to you

charleshmartin··on How to tell if you have trained your model with enough data
Great question. 4 is at the high edge of the fat tailed universality class. Most high performing models have alpha approaching 2, or at least below 3. See Figure 8(a) in the Nature paper, and our upcoming JMLR paper https://arxiv.org/abs/1810.01075
charleshmartin··on Are deep neural networks dramatically overfitted? (2019)
You can check for some signatures of over-fitting using the weightwatcher tool

https://calculatedcontent.com/2021/04/04/are-your-models-ove...

The tool identifies weight matrices that display atypical behavior, where the correlation is concentrated about unusually large matrix elements.

The idea comes from statistical mechanics of generalization, where it is known that neural networks that are over-fit are atypical and are in the spin glass phase of the learning phase space.

charleshmartin··on So You Want to Learn Physics (2016)
Agreed
charleshmartin··on So You Want to Learn Physics (2016)
Learning physics is not the same as learning to do physics. If you want to learn how to really do it, you need to work problems.

This means you to learn the techniques, and you need problem books, with solved examples, that you can work through on your own, or with a small group. This is absolutely critical. This includes both complex mathematical calculations as well as back-of-the-envelope calculations.

Moreover, while textbooks are great references, nothing replaces a good video lecture or presentation, that gets to the heart of the physical concepts without being overly technical. n And remember that physics is fundamentally an experimental science. It is not just advanced mathematics. Even if you yourself are not an experimentalist, you need to understand the basics.

charleshmartin··on Ask HN: Freelancer? Seeking freelancer? (April 2020)
Available. Expertise in ruby, python, ML, AI, data science
charleshmartin··on Ask HN: What projects are you working on now?
I'm developing the weightwatcher tool for deep neural networks into a full fledged product

http://github.com/CalculatedContent/WeightWatcher

The weightwatcher lets you detect potential problems in a trained neural network

https://calculatedcontent.com/2020/02/16/weightwatcher-empir...

charleshmartin··on Relativistic Quantum Chemistry
How it is done: https://aip.scitation.org/doi/10.1063/1.1906206
charleshmartin··on Gradient Descent Finds Global Minima of Deep Neural Networks
I would say

“Understanding Deep Learning Requires REMEMBERING Generalization”

https://calculatedcontent.com/2018/04/01/rethinking-or-remem...

https://arxiv.org/abs/1710.09553

We can understand this using the traditional theory of Statistical Mechanics of Generalization

Briefly, shuffling the labels corresponds to decreasing the effective load on the Neural Network, which pushes the system into the spin glass phase

charleshmartin··on Gradient Descent Finds Global Minima of Deep Neural Networks
An excellent paper which uses (some of the) results we have also found studying the weight matrices of neural networks...namely that they rarely undergo rank collapse

https://calculatedcontent.com/2018/09/21/rank-collapse-in-de...

But they miss something..the weight matrices also display power law behavior.

https://calculatedcontent.com/2018/09/09/power-laws-in-deep-...

This is also important because it was suggested in the early 90s that Heavy Tailed Spin Glasses would have a single local mimima.

This fact is the basis of my early suggestion that DNNs would exhibit a spin funnel

charleshmartin··on Exact mapping between Variational Renormalization Group and Deep Learning (2014)
The paper is pretty simple. It just shows that Hinton's scheme for introducing hidden variables, and then minimizing, is a type of variational RG, introduced by Kadanoff 10 years earlier

Beyond that, the analogy is not very strong, IMHO. But it is cool.

charleshmartin··on Exact mapping between Variational Renormalization Group and Deep Learning (2014)
See my blog: https://charlesmartin14.wordpress.com/2015/04/01/why-deep-le...

I try to highlight the paper and some of the history and relevance, in a general way, but also hitting the math hard.

the more general idea, which I am still formulating, is that DL systems are very different from traditional ML (ala VC theory)

In traditional ML systems, we tune the regularizer to optimize the capacity of the learner

In deep learning, we would optimize both the capacity (entropy) of the learner, and the optimization problem (energy function)

This is also what happens in the stat mech of protein folding, where the energy is optimized, even when we are at minimum capacity. This gives rise to a funneled energy landscape

https://charlesmartin14.wordpress.com/2015/03/25/why-does-de...

Similar behavior is seen generally in stat mech near a critical point, and this is why the RG analogy is relevant for me.

We should be able to see this behavior if we simply plot the entropy vs energy of say an RBM. Im not entirely sure yet how general this is, but it works for MNIST.

I discuss (some of) this in more detail in a video from a talk I gave this summer at MMDS https://www.youtube.com/watch?v=kIbKHIPbxiU

charleshmartin··on Ask HN: Freelancer? Seeking freelancer? (June 2016)
SEEKING WORK Location: San Francisco, CA

Experts in machine learning, data science, and software development.

NLP, Ad click prediction, anomaly detection, image and signal classification, deep learning.

Java, Python, Ruby.

I have worked with eBay, Blackrock, Aardvark (now Google), eHow (Demand Media), GoDaddy, ...

http://calculationconsulting.com

charleshmartin··on Deep Learning Without Poor Local Minima
That is helpful thanks.
charleshmartin··on Deep Learning Without Poor Local Minima
It is difficult to understand the implications of these assumptions and if they really apply to supervised deep learning nets.

It has been know for a very long time that simple models, like the Random Energy Model, displays a spin glass transition at low temperature. (the REM is a p-infinite limit the mean field p-spherical spin glass used in the earlier papers by LeCun) So it is expected that a random network may also display this kind of behavior, and, therefore, there could exist a large number of global minima, separated by very large barriers.

see, for example http://guava.physics.uiuc.edu/~nigel/courses/563/essays2000/...

However, it has been argued that this is unphysical for very strongly correlated systems like proteins and random hetero-polymers (with non-local contacts). Instead, one would have a spin glass minimal frustration. This would lead to a highly funneled energy landscape, with a single (or a few local) minima. This would lead to a rugged convexity.

https://charlesmartin14.wordpress.com/2015/03/25/why-does-de...

It would be helpful if some practitioners could comment on the reliability of the assumptions in this paper.

charleshmartin··on Why Deep Learning Works II: the Renormalization Group
I will address the supervised vs unsupervised issue in my next post. Here, I believe the analogy would be that when a field is applied to a spin glass, it does not exhibit a glass transition to a non-self-averaging (highly non-convex) ground state.

As to supervised vs reinforcement learning, its not that different. See how Vowpal Wabbit incoporates both the 2 ideas in how the SGD update is formulated.

charleshmartin··on Why Deep Learning Works II: the Renormalization Group
It is easier to apply the Sornette theory to antibubbles.

Bitcoin seemed like a great example.

I gotta go back and see how well the predictions actually worked.

Page 1 of 2Next →