Perhaps I'm wrong, but it seems to me that deciding on the bucket is the discrete decision. If you have two "words"/"contexts" in a sequence that ought to attend to each other, but they don't get bucketed together early in training, then there is no gradient pushing those two hidden states to be close to each other, because there is no comparison being done between the two contexts.
In a standard transformer, on the backprop we can see something like "oh, you would have been quite closer to the correct answer on this sentence if you had matched the context for 'dog' with the context for 'treat' about 20 words back." But, here, if 'dog' doesn't get bucketed with 'treat', then there's no such gradient pressure.
Eventually (and with enough hashing+bucketing), the embedding of the more relevant contexts will move closer together, but I'd suspect this might occur more slowly. Here's the authors describing the process:
> We don’t differentiate through the hash bucket assignment procedure, or the choice of what order to sort the items into. Rather, these operations take query/key vectors as input where LSH maps nearby vectors to the same bucket with high probability. Therefore, the sorting re-adjusts any time parameter updates to cause relevant vector pairs to have higher dot product, and “unhelpful” vector pairs to have lower dot products.
e: And here is a reviewer noting what I suspected about number of gradient updates,
> the performance achieved by the proposed method after 140k iterations is achieved by the full attention after ~40k iterations [on imagenet64]