As someone who spent the better part of a year trying to get various Nvidia inference products to work _at all_ even with a direct line to their developers, I will simply say "beware".
Use SGLang, vLLM, or text-generation-inference instead.
Good for the goose, good for the gander...
Can you say more?