the real gold will be when this gets finetuned. (maybe by mistral...)
Maybe its just semantics, it is technically a finetune... But to me theres a big difference between expensive "continuation training" (like Solar 10.7B or Mistral 70B) and a much less intense finetuning. The former is almost like releasing a whole new base model.
It would be awesome if Mistral did that with their data, but thats very different than releasing a Gemma Instruct finetune.
is the flow like this?
- take small dataset
- generate bigger dataset using mistral (how this is this done?)
- run LoRA to fine tune gemma extended dataset.