Ask HN: What do people mean when they say transformers aren't fully understood?
However, in online discussions about transformers -- including on this site -- I frequently see references made to some kind of mysterious, unknown element; basically, the idea that "we don't REALLY know how transformers produce such good results, at least in detail."
My question is: is that claim truly accurate? I am struggling to understand how a technology that is not fully understood, at least by specialists in the ML field, could be harnessed to such great effect. Thank you!