Looks pretty cool, congrats so far! Do you allow downloading the fine tuned model for local inference?
We utilize LoRA for smaller models, and qLoRA (quantized) for 70b+ models to improve training speeds, so when downloading model weights, what you get is the weights & adapter_config.json. Should work with llama.cpp!