This is a personal impressionistic opinion, so please take it as such and don’t expect citations or details. It has been empirically clear since the inception of GNN that the message passing structure and forced equivariance would make it hard to optimize GNNs for many practical applications. In a sense if you could solve graph isomorphism with these architectures, you would find a counter example for an NP-complete problem, so few people expected magic, but it was still disheartening to see the old school hashed bags of subgraphs outperform fancy GNNs on practical problems and to see that pretraining didn’t help the GNNs as much, or to see that the internal structures of these models are not quickly building up sensible hierarchies on their own. So then people started the hacking and exploring and by now the practical models are better than the early hacks but the theoretical optimization for how to deal with graphs in subdomains of extreme interest (say drug discovery, biology, or materials design) has not yet finished. By comparison autoregressive language models are super mature and generally applicable, as are transformers. These mixups in the early GNN led people to spend a lot of effort reformulating problems to match established architectures better or complement the GNN parts with hacks.