Have you documented your VSCode setup somewhere? I've been looking to implement something like that. Does your setup provide next edit suggestions too?
Of course, the market segment who would be most interested, probably has the expertise and funds to setup something with better horsepower than could be offered in a one size fits all solution.
Still, glad to see someone is making the product.
But if you follow the podman instructions for cuda, the llama.cpp shows you how to use their plugin here