I ask because I’ve seen posts from NVIDIA support saying that divergence is still very expensive and I’ve also seen benchmarks that force divergence in each warp by evenly splitting the warp, and the benchmarks result in 2x runtime when that happens vs. when the control-flow is dynamically uniform.
One thing to keep in mind is that even if you were to dynamically reform-warps, there’s still a potential expense because you then lose the advantage of doing things like accessing adjacent elements of data in adjacent threads. You’re bound to now have more bank conflicts, fewer memory accesses being coalesced, etc. Perhaps they do actually do this warp re-formation, but that itself does have this additional cost.