Graph Neural Networks – An Overview
theaisummer.com
theaisummer.com
It's an area I've recently been researching and they do seem to be gaining a significant amount of traction. If anyone is interested in additional reading material, I can suggest the very recent GNNs: Models and Applications (slide deck available on the website) [0].
There is also a fairly comprehensive GitHub repo on [1], though I personally haven't given it a detailed look yet.
[0] http://cse.msu.edu/~mayao4/tutorials/aaai2020/
[1] https://github.com/benedekrozemberczki/awesome-graph-classif...
They're also not very efficient, because every "feed forward step" you would have to use a matrix that is N*N (where N is the number of neurons), which is a worst case scenario. Maybe with sparse matrices it could be reasonably efficient if most weights were zero. I don't think sparse matrices are used much in machine learning currently.
These are my thoughts as a machine learning novice.
Any DAG you choose is equivalent to a network with some number of dense layers with some of the weights zeroed, so you aren't losing any modeling capability by sticking with dense networks. The current trend of massively overparameterizing networks and training them for a long time in the zero error regime (so-called "double dip" error) exploits this idea a little. With sufficient regularization, you bias the network toward a _simple_ explanation -- one where most weights are near zero, effectively the same as having chosen an optimal topology from the beginning.
If you're talking about cyclic directed graphs, those are implemented in places too, but they're extremely finicky to get right. You start having to worry about a time component to any signal propagation, you have to worry about feedback loops and unbounded signals, they're harder to get to converge during training, and so on. Afaik there isn't a solid theoretical reason why you might want to add cycles since the layered approach can already handle arbitrary problems (not that we shouldn't keep investigating -- I'm sure some people know more than me on the topic, and I don't think there's any definitive proof that cycles are always worse either, so it seems like it might be worth investigating even from a practical point of view).
Isn't it just that backpropagation on the layered topology is relatively straightforward?
That's not to say you can't write a backpropagation on an arbitrary digraph, but as you get to more and more complex digraphs, things will get harder.
I could be wrong on this.
Moreover, any arbitrary digraph can be expressed as a layered topology (possibly with a lot of 0-weights). Since there's no fundamental difference you might as well work with whatever's easiest to compute with.
Also were there any new real SOTA on any NLP tasks since last summer? I feel like accuracy progress has frozen..
What I would love would be to get a notification/mail when a task from which I subscribed got a new SOTA (from paperswithcode.com obviously).
At tasks that actually involve graphs presumably. https://arxiv.org/pdf/1901.00596.pdf https://paperswithcode.com/task/graph-classification has a bunch of GNNs ranked #1
> Also were there any new real SOTA on any NLP tasks since last summer? I feel like accuracy progress has frozen..
That's pretty normal for winter/spring, it's not conference season.
Just in the last week there are two papers which claim SOTA on different tasks.
Microsoft released Turing NLG (https://www.microsoft.com/en-us/research/blog/turing-nlg-a-1...) recently which claims SOTA on a couple tasks, it seems like the same transformer architecture, but with more layers and parameters, made feasible by training efficiency improvements. The biggest one seems to be partitioning how the model learns across different processes instead of replicating those states, which significantly improves communication, memory overhead, and training speed.
Deepmind released the Compressive Transformer (https://deepmind.com/blog/article/A_new_model_and_dataset_fo...) which claims SOTA on two other "long-range" benchmarks. My understanding of the improvement here is that instead of discarding older states, in the traditional attention layer, the compressive transformer learns which states to keep, and which states to remove.
I think these are two good examples of paper archetypes -- one where SOTA is achieved through more layers/training data/neurons (and the more interesting contribution is the improvements to model parallelism in training), and one where SOTA is achieved through a new/improved model architecture.
I wonder, for most industrial practitioners, how much either paper is useful though. The Microsoft paper helps for training billion parameter models, but most won't train a model that deep; the Deepmind paper helps for training models over very long sequences, but most people aren't using book-length sequences.*
* I remember reading somewhere that the attention mechanisms tend to only remember around 5 states (would love to see a source or study on this), which is pretty short, so would be interesting to try this model and see if the compressed transformer/attention mechanism can remember longer sequences.
I was following the work of https://www.octavian.ai/, but have not seen much recently. http://web.stanford.edu/class/cs224w/info.html also looks interesting. looking briefly over the student projects, there appear to be a couple nlp projects (e.g. http://web.stanford.edu/class/cs224w/project/26418192.pdf) but most leverage knowledge graph concepts.
I'm not sure what kind of problems you are trying to use them for, but generally they are really useful for node classification or recommendation type tasks.
So for example in the knowledge graph context, you can do things like give it a tiger, jaguar and panther and it will find nodes like lions and leopards.
https://arxiv.org/pdf/1901.00596.pdf is a survey paper which has a decent overview of the tasks they are useful for. http://research.baidu.com/Public/uploads/5c1c9a58317b3.pdf is using them for QA, but I'm not familiar enough with the dataset to evaluate it fairly.
The papers you linked are very interesting, I will have to dig further! One more recent writeup: https://eng.uber.com/uber-eats-graph-learning/ -- a real production use case, seems promising to explore more.
On the other hand, larger and larger transformer networks are constantly improving.
Assuming you are in the Northern Hemisphere so "last summer" means Julyish, then Google's T5[1] and Microsoft Turing-NLG[2] come to mind.
I find keeping an eye on HuggingFace's models list[3] is useful for this.
[1] https://arxiv.org/abs/1910.10683
[2] https://www.microsoft.com/en-us/research/blog/turing-nlg-a-1...
[3] https://github.com/huggingface/transformers#model-architectu...