If that's right, it means that in practice the proposed parallelization method will likely be much slower and much less efficient than modern implementations of self-attention, which have O(n²) time complexity and O(n) space complexity (for example, with FlashAttention). Ouch.
og_kalu, have you had a chance to look at this closely or tinker with it?