The result is identical to regular attention in transformers but training can be about four times faster, so there is almost no reason to not use it.
Not quite. There can be non-deterministic race conditions, and strange head size and sequence length requirements.
Yes. For a model within the limits of the head requirements, however, you wouldn’t be able to see a quality difference from regular attention. Non determinism is a performance price; regular transformers may also suffer from it depending on the implementation.