You can write the cuda kernels and directly differentiate them using Enzyme. Last time I checked KOKKOS was still WIP: https://github.com/EnzymeAD/Enzyme/issues?q=kokkos
5 karma · joined December 3, 2022
e.g.:
https://github.com/openxla/xla/issues/33092 https://github.com/openxla/xla/issues/35556
Explanation from the gopjrt dev:
Eventually we went with pytorch only support for the time being, with still exploring OpenXLA in place of ONNX, as a universal adapter: https://github.com/ipcamit/colabfit-model-driver