Do concurrent loops can also be parallelized.
So it is possible for a conforming program to be non-parallelizable, due to holes in the default data localization rules, despite the name of the construct and the obvious intent of the long list of restrictions imposed on code in the construct.
I summarized the two specific problems in https://github.com/llvm/llvm-project/blob/main/flang/docs/Do....
Also, isn't your employer promoting do concurrent as a method of GPU parallelization? Has this been controversial within Nvidia?