Graph NN layer: f(X, A) = g(AXW), where g is an activation
Fully connected layer: f(X) = g(XW), like GNN where the adj. matrix A=I
To put it in context a few years ago GNNs became a hot field. They are very similar in a way to transformers because both do pairwise interactions between elements, the difference being that GNNs use explicit and transformers implicit graphs.
EDIT: found it! https://arxiv.org/pdf/1609.02907.pdf
[1] : https://nanonets.com/blog/information-extraction-graph-convo...
The early graphsage stuff was, afaict, proven for generic social recommendors, but most gnn's I see seem pretty custom (e.g., deepmind's protein folding solution), esp. when not prohibitively slow. It sounds like more generic use is becoming practical w/ these libs, and esp. interesting to me, the latest NIPS had graph transformers papers, which brings another level of practically here. Not sure if DGL & friends have those yet..