Sure. The performance-purist in me would be very doubtful about the result's optimality, though.
Everything is an abstraction and choosing the right level of abstraction for your usecase is a tradeoff between your engineering capacities and your performance needs.
During the build, build.rs uses rustc_codegen_nvvm to compile the GPU kernel to PTX.
The resulting PTX is embedded into the CPU binary as static data.
The host code is compiled normally.https://github.com/Rust-GPU/Rust-CUDA/blob/aa7e61512788cc702...
The fact that different hardware has different features is a good thing.
In any case, ideally, the level of abstraction would be higher, with little application logic requiring GPU architecture awareness.