Implementing Weight-Decomposed Low-Rank Adaptation (DoRA) from Scratch
magazine.sebastianraschka.com
magazine.sebastianraschka.com
Not related to the article but tangentially relevant, would it be possible to train a LoRA or DoRA with a high rank, and then use SVD to see if the rank is too high and truncate to a better value of r? Maybe use different ranks for different layers after some training?
I haven't tried what you were suggesting, but that sounds actually plausible. Interesting idea!
"Improving LoRA: Implementing Weight-Decomposed Low-Rank Adaptation (DoRA) from Scratch"
(If it's too long, just drop the "Improving LoRA: " part)