If you have any questions about the models or training be sure to reach out and ask! I will try to respond promptly. https://twitter.com/EnricoShippole/status/165559930145459404...
The models are also compatible with many of Lucidrain's popular repositories such as Toolformer-pytorch, PaLM-rlhf-pytorch, and PaLM-pytorch. Please be sure to sponsor and help support Phil's great work: https://github.com/lucidrains/PaLM-rlhf-pytorch
You can find the weights on Hugging Face if you prefer to download the PyTorch .pt files from there instead: https://huggingface.co/conceptofmind/palm-1b
All of the C4 data has been pre-tokenized with the GPTNEOX tokenizer and blocked at sequence lengths of 8192. This will help to save you the large cost of preprocessing data. The datasets are available on Hugging Face. An example chunk can be found here: https://huggingface.co/datasets/conceptofmind/c4_0-to-20_neo...
The original unprocessed C4 dataset used can be found on the Hugging Face hub here: https://huggingface.co/datasets/c4
If you would like to preprocess your own dataset for training there is a dataset builder script provided. This uses Hugging Face datasets to efficiently map, tokenize, and block the data: https://github.com/conceptofmind/PaLM/blob/main/build_datase...
A distributed training script is provided so that you may train or fine-tune your own PaLM models using Hugging Face accelerate. More information and experiments about the training will be detailed in the repository: https://github.com/conceptofmind/PaLM/blob/main/train_distri...
The models were trained with Flash Attention, Xpos Rotary Embeddings for better length extrapolation, and multi-query single-key-value attention for more efficient decoding: https://arxiv.org/abs/2205.14135
Further instruction-tuning will be done on the new FLAN datasets we have released. https://huggingface.co/datasets/conceptofmind/cot_submix_ori...
You can find his twitter here: https://twitter.com/ShayneRedford
A basic inference script was provided in the repository which you can play around with. You may want to experiment with hyperparameters in order to get generations of varying quality. Changing a variable such as temperature matters a lot: https://github.com/conceptofmind/PaLM/blob/main/inference.py
Different inference optimizations such as Flash Attention, Hidet, and Torch compile are used. You can read more about the Hidet compiler and project here: https://pytorch.org/blog/introducing-hidet/
Our work on Toolformer, PaLM, and related projects is all thanks to the generous sponsorship by CarperAI and StabilityAI: https://github.com/CarperAI
This is not an official Google or StabilityAI product.