Transformers Are Graph Neural Networks
graphdeeplearning.github.io
graphdeeplearning.github.io
This is the slow decline of the machine learning field because most researchers are too busy procuring positions in various institutions instead of asking and pursuing more creative and more difficult lines of question/thought.
I am glad the author chooses to call out disconcerting behavior.
> these papers showed that Transformer heads can be ‘pruned’ or removed after training without significant performance impact.
The larger the model the less important or just purely redundant the various modules will become.
Correlation may be != causation. But absent other evidence, I would still be careful with accusations of any systemic issues with their work culture, or predictions about impending doom.
I seem to remember some of these seemingly secondary parameters at one point being the sole reason making the model work. Wasn’t it a new initialization schedule that kicked of the current boom?
In any case, recent history should be a good example of how “just” uncreative pursuits such as increasing model depth can have results of dramatically different quality.
It also feels strange to take issue with people being motivated by publication. As far as inducing altruistic behavior goes, publications are second only to cheap medals handed posthumously to the children of dead soldiers. And in terms publication criteria being aligned with some abstract sense of “good research”, I have little doubt that creative ideas with some interesting results will find an interested audience. There may well remain the universal problem that unsuccessful “out-there” efforts may leave you with little when they fail, but risk is as inherent to such efforts as the chance to make it big. It’s almost tautologically impossible to adequately reward failed efforts, because there are no measures to asses them; indeed, where there are measure, they are no longer deemed to have failed.
To then make it easier to advance along new lines of thinking, we would want to come up with new yardsticks to judge results: coming up with new standard problem sets where current efforts fail dramatically. Thinking back over the last years, I’m not entirely sure that isn’t exactly what we’ve been doing.
No lesser man than Geoff Hinton himself thinks there are systemic issues with machine learning publications, although he doesn't foresee impending doom:
WIRED: The recent boom of interest and investment in AI and machine learning means there’s more funding for research than ever. Does the rapid growth of the field also bring new challenges?
GH: One big challenge the community faces is that if you want to get a paper published in machine learning now it's got to have a table in it, with all these different data sets across the top, and all these different methods along the side, and your method has to look like the best one. If it doesn’t look like that, it’s hard to get published. I don't think that's encouraging people to think about radically new ideas.
Now if you send in a paper that has a radically new idea, there's no chance in hell it will get accepted, because it's going to get some junior reviewer who doesn't understand it. Or it’s going to get a senior reviewer who's trying to review too many papers and doesn't understand it first time round and assumes it must be nonsense. Anything that makes the brain hurt is not going to get accepted. And I think that's really bad.
What we should be going for, particularly in the basic science conferences, is radically new ideas. Because we know a radically new idea in the long run is going to be much more influential than a tiny improvement. That's I think the main downside of the fact that we've got this inversion now, where you've got a few senior guys and a gazillion young guys.
https://www.wired.com/story/googles-ai-guru-computers-think-...
That's actually been part of the reason for it's explosion. See Google, Facebook, Amazon, Uber, all releasing public blurbs and repos of their research. And papers.
The multidisciplinary nature of ML attracts science types who value publications. It's a sort of club. I'm talking about the people doing really cool and difficult stuff - robots, AI, self driving, cutting edge and high latent value. It's a strong incentive for them.
"Sentences are a fully connected graph". Ok fine, but that's a graph with basically no information embedded in its structure. GNNs are supposed to be useful for graphs that have interesting structure, right?
Can you explain this?
Of course, this connection may be trivial to most people, but I hadn't seen a post on this before. So I decided to write one for myself as I studied these architectures.
Sentences can be reasonable modeled as a mostly 1 directional graph.