This is an extension of our work, Shayne Longpre and I, on an open-source reproduction of the FLAN V2 dataset. https://twitter.com/EnricoShippole/status/166175616624899686...
Due to both the nature of the training of OpenLLaMA and the curation of the FLAN dataset, we have been working to realize permissively licensed, quality, instruction-tuned models and datasets similar to the original FLAN-T5 suite.
The compute for this model release is all thanks to the generous sponsorship by @carperai, @EMostaque, and @StabilityAI.
A big thank you to @zhansheng of @AiEleuther and @fabmilo for helping build the dataset as well.
The base OpenLLaMA-7b model used in these experiments was developed by @younggeng and @haoliuhl and can be found here: https://t.co/IBySrP65d3
I previously worked with @ShayneRedford the main author of the FLAN collection to recreate his great work and publicly release high-quality instruction tuning data. https://twitter.com/ShayneRedford/status/1661734033720762368...
You can find the original FLAN repository and all of @ShayneRedford's incredible work here: https://github.com/google-research/FLAN/tree/main/flan/v2#do...
We are going to soon be releasing a massive causal language model dataset containing hundreds of GBs of high-quality instruction data. Look out for that release in the near future.
You can check out @ShayneRedford's new paper on building pre-training datasets here:https://twitter.com/ShayneRedford/status/1660670374206652419...
This is not an official @StabilityAI product.
If you have any questions about the data or model be sure to reach out and ask! I will try to respond promptly.
A Flan-Open-Llama-13b model will be released soon.