How does the attention operator in transformers, in which input data is multiplied by input data (as opposed other neural network operations in which input data is multiplied by model weights) fit into the notion of a universal activator?
One unsatisfying argument might be that this might fall into implementation details for this particular class. Another prediction might be that an attention mechanism is an essential element of these networks that appears in other networks of this class. Another is that this is a decent approximation, but has limitations, and we'll figure out how the brain does it and replace it with that.