We worked directly with Kaiokendev, to extend the context length of the open-llama 7b and 3b models through fine-tuning. The fine-tuned models maintain the same perplexity at 8k extrapolation surpassing the performance of other recent methodologies.
Applying the method to the rotary position embedding requires only slight changes to the model's code by dividing the positional index, t, by a scaling factor.
You can review the benchmarks of the LLongMA models trained at 8k when compared to the original Open-LLaMA models trained at 2k. We show slight performance boosts on multiple benchmarks and minimal degradation on others.
The LLongMA 7b model is available on huggingface to use:https://huggingface.co/conceptofmind/LLongMA-7b
The LLongMA 3b model can also be found on huggingface here:https://huggingface.co/conceptofmind/LLongMA-3b
The repository containing theemozilla’s implementation of scaled rotary embeddings can be found here:https://github.com/jquesnelle/scaled-rope
A LLongMA-13b model trained at 8k context length will be released soon. As well as a suite of LLongMA models trained at 16k and 32k context lengths.
If you would like to learn more about scaling rotary embeddings, I would strongly recommend reading Kaiokendev's blog posts on his findings:https://kaiokendev.github.io/
A PR to add scaled rotary embeddings to huggingface transformers is currently in progress: https://github.com/huggingface/transformers/pull/24653
The model was trained for ~1 billion tokens on togethercompute's Red Pajama dataset. The context length of the examples varies: https://huggingface.co/datasets/togethercomputer/RedPajama-D...
The pre-tokenized dataset will be available here for you to use soon: https://huggingface.co/datasets/conceptofmind/rp-packed-8k-n...
I would also recommend checking out the phenomenal research by OfirPress on ALiBi which laid the foundation for many of these scaling techniques: https://arxiv.org/abs/2108.12409
OfirPress just live-streamed his Ph.D. defense recently here:https://twitter.com/OfirPress/status/1677025700560388098?s=2...
It is also worth reviewing the paper, A Length-Extrapolatable Transformer, and xPos technique which also applies scaling to rotary embeddings: https://arxiv.org/pdf/2212.10554.pdf
We previously trained the first publicly available model with rotary embedding scaling here: https://twitter.com/EnricoShippole/status/165559930145459404...
The compute for this model release is all thanks to the generous sponsorship by carperai, EMostaque, and StabilityAI.
You can find out more about the NousResearch organization here:https://huggingface.co/NousResearch
The base OpenLLaMA models used in these experiments were developed by younggeng and haoliuhl and can be found here: https://huggingface.co/openlm-research
This is not an official StabilityAI or NousResearch product.
Credits to Zanzibased for the llama image.
If you have any questions about the data or models be sure to reach out and ask! I will try to respond promptly.