So this can tune a model 7X faster than LoRA, which was already a massive speed boost? Curious to see what this will do to the LLaMA-derivative community in particular.
I am not convinced that the "best rank" is not just the highest possible with your compute budget, personally.