This is seriously impressive for a single project. Rust + CUDA kernels for MD, fallback to SIMD thread pools, pharmacokinetics inference, docking, and it compiles to a standalone binary? How does the MD throughput compare to GROMACS on equivalent hardware?