A typical LLM might use about 0.1% of CUDA. That's all that would have to be ported to get that LLM to work.
Then again, maybe the goal is getting 0.1% of CUDA market share. /s
You are mostly listing irrelevant nice to have things that aren't deal breakers. AMD's consumer GPUs have a long history of being abandoned a year or two after release.
Coupled with Khronos, Intel, AMD never delivering anything comparable with OpenCL, Apple losing interest after Khronos didn't took OpenCL into the direction they wanted, Google never adopting it favouring their Renderscript dialect.