It seems we could repeatedly apply attention as an iterative update rule, feeding the output sequence as the new queries: new_Q = Attention(old_Q, K, V). Obviously, no one ever does that. But in theory we could, and apparently it would work just as well.
In self-attention, we apply dense layers to the input sequence to get K and V, so we can rewrite the update rule like this: new_Q = Attention(old_Q, input_seq). Internally, Attention would have to compute and cache K and V from input_seq. The input sequence is given and remains fixed, so we can rewrite the update rule again to make that explicit, like this: new_Q = Attention(old_Q | input_seq).
According to the preprint, everything the routing algorithm does can be written as an update rule U that looks just like that, except that the preprint calls the queries "x_out" and the input sequence "x_inp". Using their notation, the update rule looks like this (section 3.3 and appendix B): x_out <-- U(x_out | x_inp).
One more thing that seems important: In self-attention, we get the initial queries by applying a dense layer to the input sequence. In the routing algorithm, the initial queries come instead from the same update rule, U, by assuming all possible output states have equal probability. Interestingly, the code in the github repo executes only two updates by default: one to compute the initial queries, and another one to get updated queries, which are returned as the output sequence. This isn't too different from what we normally do with attention: first we compute the initial queries, and then we apply one update to get updated queries, which are returned as the output sequence.
EDITS: Many.