A modern self-referential weight matrix that learns to modify itself
arxiv.org
arxiv.org
The failure to put something like that front and center makes me wonder how strong the method is, because you have to assume that someone on the team has tried more benchmarks. Still, the idea of learning a better update rule than gradient descent is intriguing, so maybe something cool will come from this :)
Second: do Hacker News posts form a small-world network? I don't know. I don't even know if my question is well posed (it might be a meaningless question). Does the set of Hacker News articles change over time in ways that resemble annealing or self-training matrices? (likewise, I question this question, but I wonder.)
I am never sure if it is a waste of time or has some value.
If you guys had some unique ML technology that is different to what all the others do, what would you do with it?
Use the "proof is in the pudding" method:
Do something with it - preferably useful - that no one else can.
For unsupervised learning algorithms like masked models (BERT and some other Transformers), it makes sense to train in parallel with prediction. Why not?
My imagination can't wrap around using this for supervised (labeled data) learning.
> The WM of a self-referential NN, however, can keep rapidly modifying all of itself during runtime. In principle, such NNs can meta-learn to learn, and metameta-learn to meta-learn to learn, and so on, in the sense of recursive self-improvement.
Everyone who doubts is hanging everything on "in principle" being too hard. Seems ridiculous to me, a failure of imagination.
Your quote sounds like it could just as well have been from that thesis.
My 5-year-old, consumer-grade GPU does 1.5 MHz * 2300 cores, whereas the equivalent released this year does 1.7 Mhz * 8900 cores. Granted, not the best way to measure GPU performance, but it is roughly keeping pace with Moore's law, and it's going to be a better indicator of the future than Intel CPU capabilities, especially for machine learning applications.