Money and popularity are orthogonal to pathfinding that leads to breakthroughs.
Transformers came out at roughly the same time[a] and have proven to be great at... pretty much everything. They just work. Since then, most AI research money, effort, and compute has been invested to study and improve Transformers and related models, at the expense of almost everything else.
Many promising ideas, including routing, won't be seriously re-explored until and unless progress towards AGI seems to stall.
---
Some kinds of explanations which I think are at least plausible (but IDK if any evidence exists for them):
- The attention structure in transformers allows any chunk to be learned to be important for any other chunk. And pretty quickly people tended towards these being pretty deep. By comparison, the capsule + routing structure (IIUC) came with a built-in kind of sparsity (from capsules at a level in the hierarchy being independent), and because the hierarchy was meant to align with composition, it often (I think) didn't have a huge number of levels? Maybe this flexibility + depth are key?
- Related to capsules being independent, an initial design feature in capsule networks seems to have been smaller model sizes. Perhaps this was at some level just a bad thing to reach for? I think at the time, "smaller models means optimization searches over a smaller space, which is faster to converge and requires less data" was still sort of in people's heads, and I think this view is pretty much dead.
- I've heard some people argue that one of the core strengths of transformers is that they support training in a way allows for maxing out available GPUs. I think this is mostly in comparison to previous language models which were explicitly sequential. But are capsule networks less conducive to efficient training?
Promising approaches are often ignored for reasons that have little to do with their merits. For example, Hinton, Bengio, and Lecun spent much of the 1990's and all of the 2000's in the fringes of academia, unable to get much funding, because few others were interested in or saw any promise in deep learning! Similarly, Katalin Karikó lost her job and spent almost two decades in obscurity because few others were interested or saw any promise in RNA vaccines!
Now, I'm not saying routing methods will become more popular in the future. I mean, who the heck knows?
What I'm saying is that promising approaches can fall out of favor for reasons that are not intrinsic to them.