ParentFull threadantimatter15·From my experimentation it seems like there's some significant loss in accuracy running the tuned LoRa models through llama.cpp (due to bugs/differences in inference or tokenization), even aside from losses due to quantization.View on HN