Isn't that going to be a problem with no NVIDIA GPUs?
It is, unfortunately, not an apples-to-apples comparison, because on the nVidia cluster I'm running it via llama-cpp-python and a quantized 34B version, while on LUMI I'm running the official non-quantized full 70B version via the transformers library.
Long story short, I'm getting a 7.5x higher throughput from LUMI than on the nVidia cluster (which means each card is 5x faster on LUMI).
Edit: The AMD GPUs work fine because one can run Pytorch for ROCm via the pytorch-triton-rocm package.
Those super computers are extremely powerful, it might not be as energy efficient as H100s, but it does the job.