HNHacker News
TopNewBestAskShowJobs

EnricoShippole

18 karma · joined July 5, 2022

submissionscomments
EnricoShippole··on Open-Sourcing SEC Edgar on Hugging Face
Given the increasingly closed-source nature of the U.S. AI ecosystem, it is now more important than ever to push for the proliferation of open model and dataset releases. Datamule, TeraflopAI, and Daft collaborated to release 43 Billion Tokens of SEC EDGAR data.
EnricoShippole··on The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text
Large language models (LLMs) are typically trained on enormous quantities of unlicensed text, a practice that has led to scrutiny due to possible intellectual property infringement and ethical concerns. Training LLMs on openly licensed text presents a first step towards addressing these issues, but prior data collection efforts have yielded datasets too small or low-quality to produce performant LLMs. To address this gap, we collect, curate, and release the Common Pile v0.1, an eight terabyte collection of openly licensed text designed for LLM pretraining. The Common Pile comprises content from 30 sources that span diverse domains including research papers, code, books, encyclopedias, educational materials, audio transcripts, and more. Crucially, we validate our efforts by training two 7 billion parameter LLMs on text from the Common Pile: Comma v0.1-1T and Comma v0.1-2T, trained on 1 and 2 trillion tokens respectively. Both models attain competitive performance to LLMs trained on unlicensed text with similar computational budgets, such as Llama 1 and 2 7B. In addition to releasing the Common Pile v0.1 itself, we also release the code used in its creation as well as the training mixture and checkpoints for the Comma v0.1 models.
EnricoShippole··on [dead]
Introducing LLongMA, a series of OpenLLaMA models, trained at 8k context length using linear positional interpolation scaling. The model was trained in collaboration with theemozilla of NousResearch and Kaiokendev.

We worked directly with Kaiokendev, to extend the context length of the open-llama 7b and 3b models through fine-tuning. The fine-tuned models maintain the same perplexity at 8k extrapolation surpassing the performance of other recent methodologies.

Applying the method to the rotary position embedding requires only slight changes to the model's code by dividing the positional index, t, by a scaling factor.

You can review the benchmarks of the LLongMA models trained at 8k when compared to the original Open-LLaMA models trained at 2k. We show slight performance boosts on multiple benchmarks and minimal degradation on others.

The LLongMA 7b model is available on huggingface to use:https://huggingface.co/conceptofmind/LLongMA-7b

The LLongMA 3b model can also be found on huggingface here:https://huggingface.co/conceptofmind/LLongMA-3b

The repository containing theemozilla’s implementation of scaled rotary embeddings can be found here:https://github.com/jquesnelle/scaled-rope

A LLongMA-13b model trained at 8k context length will be released soon. As well as a suite of LLongMA models trained at 16k and 32k context lengths.

If you would like to learn more about scaling rotary embeddings, I would strongly recommend reading Kaiokendev's blog posts on his findings:https://kaiokendev.github.io/

A PR to add scaled rotary embeddings to huggingface transformers is currently in progress: https://github.com/huggingface/transformers/pull/24653

The model was trained for ~1 billion tokens on togethercompute's Red Pajama dataset. The context length of the examples varies: https://huggingface.co/datasets/togethercomputer/RedPajama-D...

The pre-tokenized dataset will be available here for you to use soon: https://huggingface.co/datasets/conceptofmind/rp-packed-8k-n...

I would also recommend checking out the phenomenal research by OfirPress on ALiBi which laid the foundation for many of these scaling techniques: https://arxiv.org/abs/2108.12409

OfirPress just live-streamed his Ph.D. defense recently here:https://twitter.com/OfirPress/status/1677025700560388098?s=2...

It is also worth reviewing the paper, A Length-Extrapolatable Transformer, and xPos technique which also applies scaling to rotary embeddings: https://arxiv.org/pdf/2212.10554.pdf

We previously trained the first publicly available model with rotary embedding scaling here: https://twitter.com/EnricoShippole/status/165559930145459404...

The compute for this model release is all thanks to the generous sponsorship by carperai, EMostaque, and StabilityAI.

You can find out more about the NousResearch organization here:https://huggingface.co/NousResearch

The base OpenLLaMA models used in these experiments were developed by younggeng and haoliuhl and can be found here: https://huggingface.co/openlm-research

This is not an official StabilityAI or NousResearch product.

Credits to Zanzibased for the llama image.

If you have any questions about the data or models be sure to reach out and ask! I will try to respond promptly.

EnricoShippole··on [dead]
Releasing Flan-Open-Llama-7b, an OpenLLaMA model fine-tuned on the FLAN instruction dataset. https://t.co/CLb3yiwOMH

This is an extension of our work, Shayne Longpre and I, on an open-source reproduction of the FLAN V2 dataset. https://twitter.com/EnricoShippole/status/166175616624899686...

Due to both the nature of the training of OpenLLaMA and the curation of the FLAN dataset, we have been working to realize permissively licensed, quality, instruction-tuned models and datasets similar to the original FLAN-T5 suite.

The compute for this model release is all thanks to the generous sponsorship by @carperai, @EMostaque, and @StabilityAI.

A big thank you to @zhansheng of @AiEleuther and @fabmilo for helping build the dataset as well.

The base OpenLLaMA-7b model used in these experiments was developed by @younggeng and @haoliuhl and can be found here: https://t.co/IBySrP65d3

I previously worked with @ShayneRedford the main author of the FLAN collection to recreate his great work and publicly release high-quality instruction tuning data. https://twitter.com/ShayneRedford/status/1661734033720762368...

You can find the original FLAN repository and all of @ShayneRedford's incredible work here: https://github.com/google-research/FLAN/tree/main/flan/v2#do...

We are going to soon be releasing a massive causal language model dataset containing hundreds of GBs of high-quality instruction data. Look out for that release in the near future.

You can check out @ShayneRedford's new paper on building pre-training datasets here:https://twitter.com/ShayneRedford/status/1660670374206652419...

This is not an official @StabilityAI product.

If you have any questions about the data or model be sure to reach out and ask! I will try to respond promptly.

A Flan-Open-Llama-13b model will be released soon.

EnricoShippole··on [dead]
With Reddit and many other sites shutting down access to their APIs it is now more important than ever to release quality open-source conversational data. I worked with the author of FLAN, Shayne Longpre, to generate ~80GB of labeled FLAN dialog data. You can find the dataset here: https://huggingface.co/datasets/conceptofmind/flan_dialog_su...

This is a previous extension of his work to publicly release the FLAN collection: https://twitter.com/EnricoShippole/status/166175616624899686...

The dialog data is available on huggingface to download. It was processed at an extended context length of 8192. It contains relevant metadata such as Inputs, Targets, Task Source, and Task Name.

An additional FLAN Dialog submix dataset was also preprocessed for causal language modeling, fixing different encoding issues, and is available on huggingface to download: https://huggingface.co/datasets/conceptofmind/flan_dialog_su...

You can find the extended FLAN Zero-shot-only causal language modeling dataset on huggingface here: https://huggingface.co/datasets/conceptofmind/flan_dialog_zs...

You can find the extended FLAN Few-shot-only causal language modeling dataset on huggingface here:https://huggingface.co/datasets/conceptofmind/flan_dialog_fs...

Our work on an open reproduction of FLAN V2 and this conversational data release is all thanks to the generous sponsorship by CarperAI and StabilityAI.

You can find more information about CarperAI here: https://carper.ai/

And StabilityAI here: https://stability.ai/

A big thank you to Jason Phang from EleutherAI for helping build the dataset as well. You can find his Twitter here: https://twitter.com/zhansheng

You can find out more about Shayne Longrpe's work on FLAN here:https://twitter.com/ShayneRedford/status/1661734033720762368

You can find the original FLAN repository here: https://github.com/google-research/FLAN/tree/main/flan/v2#do...

Be sure to also read through Shayne Longpre's paper on the FLAN collection to get better insight into how the data was created:https://arxiv.org/abs/2301.13688

You can check out Shayne Longpre's new paper on building pre-training datasets here:https://twitter.com/ShayneRedford/status/1660670374206652419

This is not an official Google or StabilityAI product.

If you have any questions about the data be sure to reach out and ask! I will try to respond promptly.

EnricoShippole··on [dead]
Open-source reproduction of the FLAN V2 dataset.

The full dataset can be found here: https://huggingface.co/datasets/conceptofmind/FLAN_2022

I worked with Shayne Longpre the main author of the FLAN collection to recreate his great work and publicly release high-quality instruction tuning data. We fixed encoding issues and also increased the sequence length to 4096: https://twitter.com/EnricoShippole/status/166175616624899686...

Each of the individual submixes is also available on huggingface to download. The sub-mixes are T0, FLAN2021, CoT, NIv2, and Dialog. Each contains relevant metadata such as Inputs, Targets, Task Source, Task Name, and Template Type.

T0 submix: https://huggingface.co/datasets/conceptofmind/t0_submix_orig... Flan2021 submix: https://huggingface.co/datasets/conceptofmind/flan2021_submi... CoT submix: https://huggingface.co/datasets/conceptofmind/cot_submix_ori... NIv2 submix: https://huggingface.co/datasets/conceptofmind/niv2_submix_or... Dialog submix: https://huggingface.co/datasets/conceptofmind/dialog_submix_...

You can find the original FLAN repository and all of Shayne Longpre's incredible work here: https://github.com/google-research/FLAN/tree/main/flan/v2#do...

Be sure to also read through Shayne's paper on the FLAN collection to get better insight into how the data was created: https://arxiv.org/abs/2301.13688

We are going to soon be releasing a massive causal language modeling dataset containing hundreds of GBs of high-quality instruction data. Look out for that release in the near future.

Our work on an open reproduction of FLAN V2 and related projects is all thanks to the generous sponsorship by CarperAI and StabilityAI. You can find out more about CarperAI here: https://carper.ai/ And StabilityAI here: https://stability.ai/

A big thank you to Jason Phang and Fabrizio Milo for helping build the dataset as well. You can find Jason Phang's twitter here: https://twitter.com/zhansheng And Fabrizio Milo's here: https://twitter.com/fabmilo

You can check out Shayne's new paper on building pre-training datasets here: https://github.com/shayne-longpre/a-pretrainers-guide/blob/m...

This is not an official Google or StabilityAI product.

If you have any questions about the data be sure to reach out and ask! I will try to respond promptly: https://twitter.com/EnricoShippole

EnricoShippole··on [dead]
Introducing three new open-source PaLM models trained at a context length of 8k on C4. Open-sourcing LLMs is a necessity for the fair and equitable democratization of AI. The models of sizes 150m, 410m, and 1b are available to download and use here: https://github.com/conceptofmind/PaLM

If you have any questions about the models or training be sure to reach out and ask! I will try to respond promptly. https://twitter.com/EnricoShippole/status/165559930145459404...

The models are also compatible with many of Lucidrain's popular repositories such as Toolformer-pytorch, PaLM-rlhf-pytorch, and PaLM-pytorch. Please be sure to sponsor and help support Phil's great work: https://github.com/lucidrains/PaLM-rlhf-pytorch

You can find the weights on Hugging Face if you prefer to download the PyTorch .pt files from there instead: https://huggingface.co/conceptofmind/palm-1b

All of the C4 data has been pre-tokenized with the GPTNEOX tokenizer and blocked at sequence lengths of 8192. This will help to save you the large cost of preprocessing data. The datasets are available on Hugging Face. An example chunk can be found here: https://huggingface.co/datasets/conceptofmind/c4_0-to-20_neo...

The original unprocessed C4 dataset used can be found on the Hugging Face hub here: https://huggingface.co/datasets/c4

If you would like to preprocess your own dataset for training there is a dataset builder script provided. This uses Hugging Face datasets to efficiently map, tokenize, and block the data: https://github.com/conceptofmind/PaLM/blob/main/build_datase...

A distributed training script is provided so that you may train or fine-tune your own PaLM models using Hugging Face accelerate. More information and experiments about the training will be detailed in the repository: https://github.com/conceptofmind/PaLM/blob/main/train_distri...

The models were trained with Flash Attention, Xpos Rotary Embeddings for better length extrapolation, and multi-query single-key-value attention for more efficient decoding: https://arxiv.org/abs/2205.14135

Further instruction-tuning will be done on the new FLAN datasets we have released. https://huggingface.co/datasets/conceptofmind/cot_submix_ori...

You can find his twitter here: https://twitter.com/ShayneRedford

A basic inference script was provided in the repository which you can play around with. You may want to experiment with hyperparameters in order to get generations of varying quality. Changing a variable such as temperature matters a lot: https://github.com/conceptofmind/PaLM/blob/main/inference.py

Different inference optimizations such as Flash Attention, Hidet, and Torch compile are used. You can read more about the Hidet compiler and project here: https://pytorch.org/blog/introducing-hidet/

Our work on Toolformer, PaLM, and related projects is all thanks to the generous sponsorship by CarperAI and StabilityAI: https://github.com/CarperAI

This is not an official Google or StabilityAI product.

EnricoShippole··on [dead]
An open-source implementation of the CrossFormer: A Versatile Vision Transformer Hinging on Cross-scale Attention research paper in Google's JAX and Flax.

JAX/Flax Repository: https://github.com/conceptofmind/Crossformer-flax

Crossformer Research Paper: https://arxiv.org/abs/2108.00154

Official Github Repository: https://github.com/cheerss/CrossFormer

In collaboration with Dr. Phil 'Lucid' Wang: https://github.com/lucidrains.

EnricoShippole··on [dead]
An open-source implementation of the CvT: Introducing Convolutions to Vision Transformers research paper in Google's JAX and Flax.

JAX/Flax Repository: https://github.com/conceptofmind/CvT-flax

CvT Research Paper: https://arxiv.org/abs/2103.15808

Official Github Repository: https://github.com/microsoft/CvT

In collaboration with Dr. Phil 'Lucid' Wang: https://github.com/lucidrains.

EnricoShippole··on [dead]
An #opensource implementation of the Twins: Revisiting the Design of Spatial Attention in Vision Transformers research paper in Google's #JAX and #Flax.

JAX/Flax Repository: https://github.com/conceptofmind/Twins-SVT-Flax

Twins SVT Research Paper: https://arxiv.org/abs/2104.13840

Official Github Repository: https://github.com/Meituan-AutoML/Twins

In collaboration with Dr. Phil 'Lucid' Wang: https://github.com/lucidrains.

EnricoShippole··on [dead]
An open-source implementation of the Rethinking Spatial Dimensions of Vision Transformers research paper in Google's #JAX and #Flax.

JAX/Flax Repository: https://github.com/conceptofmind/PiT-flax

PiT Research Paper: https://arxiv.org/abs/2103.16302

Official Github Repository: https://github.com/naver-ai/pit

In collaboration with Dr. Phil 'Lucid' Wang: https://github.com/lucidrains.

EnricoShippole··on Open-Source ScalableViT Implementation
An open-source implementation of the ScalableViT: Rethinking the Context-oriented Generalization of Vision Transformer research paper in Google's JAX and Flax.

"The vanilla self-attention mechanism inherently relies on pre-defined and steadfast computational dimensions. Such inflexibility restricts it from possessing context-oriented generalization that can bring more contextual cues and global representations. To mitigate this issue, we propose a Scalable Self-Attention (SSA) mechanism that leverages two scaling factors to release dimensions of query, key, and value matrix while unbinding them with the input. This scalability fetches context-oriented generalization and enhances object sensitivity, which pushes the whole network into a more effective trade-off state between accuracy and cost. Furthermore, we propose an Interactive Window-based Self-Attention (IWSA), which establishes interaction between non-overlapping regions by re-merging independent value tokens and aggregating spatial information from adjacent windows. By stacking the SSA and IWSA alternately, the Scalable Vision Transformer (ScalableViT) achieves state-of-the-art performance in general-purpose vision tasks. For example, ScalableViT-S outperforms Twins-SVT-S by 1.4% and Swin-T by 1.8% on ImageNet-1K classification." - Rui Yang, Hailong Ma, Jie Wu, Yansong Tang, Xuefeng Xiao, Min Zheng, Xiu Li

ScalableViT Research Paper: https://arxiv.org/abs/2203.10790

In collaboration with Dr. Phil 'Lucid' Wang: https://github.com/lucidrains.

Developer updates can be found on: https://twitter.com/EnricoShippole https://www.linkedin.com/in/enrico-shippole-495521b8/

EnricoShippole··on Open-Source Simple-ViT Implementation
An open-source implementation of the Better plain ViT baselines for ImageNet-1k research paper in Google's JAX and Flax.

An update from some of the same authors of the original paper proposes simplifications to ViT that allows it to train faster and better.

Among these simplifications include 2d sinusoidal positional embedding, global average pooling (no CLS token), no dropout, batch sizes of 1024 rather than 4096, and use of RandAugment and MixUp augmentations. They also show that a simple linear at the end is not significantly worse than the original MLP head.

Simple ViT Research Paper: https://arxiv.org/abs/2205.01580

Official Github repository: https://github.com/google-research/big_vision

Developer updates can be found on: https://twitter.com/EnricoShippole

In collaboration with Dr. Phil 'Lucid' Wang: https://github.com/lucidrains

EnricoShippole··on Open-Source CaiT Implementation
An open-source implementation of the Going deeper with Image Transformers research paper in Google's JAX and Flax.

The paper also notes difficulty in training vision transformers at greater depths and proposes two solutions. First it proposes to do per-channel multiplication of the output of the residual block. Second, it proposes to have the patches attend to one another, and only allow the CLS token to attend to the patches in the last few layers.

CaiT Research Paper: https://arxiv.org/abs/2103.17239

Official Github repository: https://github.com/rwightman/pytorch-image-models

Developer updates can be found on: https://twitter.com/EnricoShippole