Show HN: Fine-tuned coding copilot with Gemma and qLORA
github.com
github.com
Curious how does the double quantization of weights in qLORA impact the model's performance compared to standard fine-tuning. Could qLORA and LoRA be combined with other optimization methods, such as pruning or knowledge distillation, for more efficiency gains?