RedPajama v2 Open Dataset with 30T Tokens for Training LLMs
together.ai
together.ai
- example summary, for better topic embedding
- RAG based summary, to have the model critically assess its training data distribution and answer questions on it; to bring together information sitting in separate examples
- named entities, for knowledge base; maybe it helps with fact checking later
- implicit tasks present in the text, what are the tasks a LLM could learn from a given example?
- chain-of-thought augmentation, to bring out implicit deductions and reduce information fragmentation; it has been shown in the Phi-1.5 paper and Orca that synthetic CoT datasets are superior source materials
What data fragmentation? Look at the Reversal Curse paper. Models that train on "A is the father of B" fail to generate "B is the son of A". This kind of connection needs to be explicitly added, and would improve task solving as well.
Training on purely organic data is not good enough anymore. All powerful models train on a mix of organic and synthetic data, some models on 50-50 proportions, like the web+synth variant from Phi-1.5.
The main idea is to go deeper into the raw data, to infuse it with insight. LLM dataset preprocessing is going to be expensive, comparable to training costs, but the results are worth the effort.
If you are interested in contributing the code for these features, feel free to do a PR to https://github.com/togethercomputer/RedPajama-Data! Otherwise we will try our best effort implementation :) but we hope that this can become a community effort
(feel free to created more issues on github for us to keep track. I created one for this https://github.com/togethercomputer/RedPajama-Data/issues/76)
B could be A's daughter.
If you take the training data, or in general, imagine taking all that humans ever wrote or spoke in English to date, you'll expect to find an overwhelming amount of cases where "The color of the sky is" ends with "blue". However, "Blue is the color of" can easily have a hundred thousand different plausible completions, and "the sky" won't even be one of the more likely ones. In the absence of additional context that strongly hints at the answer, one should NOT expect a properly working LLM to frequently propose "the sky" as completion to "Blue is the color of".
To further elaborate on this (possibly wrong) understanding:
- Each person can then run their own processing, possibly duplicating effort(?)
...but on the good side, giving each person the ability to tweak the pipeline to suit their needs.
- There is no torrent of already processed data because __________?
- Looking at file lists for this on Hugging Face, some files seem to be stored in Git Large File Storage. Are these already processed files that together constitute the dataset? Or are these Common Crawl files that are selectively listed and pulled for processing?
What options are there to preemptively obtain a copy, in case of any possible eventual takedown of the dataset, any assurances about access aside? I am reminded of parts of the pile.
Obviously I'm super clueless here... please be gentle and share anything you know or correct anything I've got wrong.
I'm not asking about training, if that wasn't obvious. Just about obtaining the dataset.
--
(A) the dataset after pre-processing the raw CommonCrawl data (e.g., text extraction and language identification) and some minimal filtering; and
(B) for each document in (A), we also pre-computed 40+ of "features" (we call the "quality annotations") you can use to further filter it or deduplicate it. For example, one such feature is "how similar this document is to Wikipedia".
--
(A) is around 30T tokens, but you might want to use features in (B) to further filter/dedup it down, e.g., to 5T. For example, if in your application documents similar to Wikipedia are the most helpful documents, you can take the top documents with the highest score for the feature "how similar this document is to Wikipedia". Of course, the really interesting case happens when you consider a larger subset of these features (or maybe even automatically learn what the best way of filtering it is).
Our goal is to make this as flexible as possible such that you can fit this into your own application. What we have released is both (A) and (B)
If you have any questions, please let us know! Thanks for your interests, have fun with the data!
> how similar this document is to Wikipedia
So that’s a measure of how similar it is to the background vector of all (language in focus) Wikipedia data?
- `rps_doc_ml_wikiref_score`: a classifier that classifiers random webpage with Wiki references (used in Llama-1)
- `ccnet_perplexity`: perplexity of an LM trained on Wikipedia (used in CCNet)
- `rps_doc_ml_wikipedia_score`: classifier prediction for the document being a Wikipedia article
- `rps_doc_wikipedia_importance`: Used in https://arxiv.org/abs/2302.03169
You can see the full table here: https://together.ai/blog/redpajama-data-v2
They state the 1 trillion token dataset is 5TB.
Is it safe to assume this is 5TB * 30 = 150TB?
The code in the HuggingFace repo downloads data from url base: https://data.together.xyz/redpajama-data-v2/v1.0.0
https://huggingface.co/datasets/togethercomputer/RedPajama-D...
https://espadrine.github.io/blog/posts/chinchilla-s-death.ht...
I think by the time the former is an "issue", we'll have a Super Intelligence on our hands anyway.
The latter is looking less and less likely to be a real hurdle. Very little inductive bias to steer away from crucial solutions, very scalable.
I don't think we're even close. Libgen's nonfiction archive alone is over 32 terabytes. Total size last year was over 120 terabytes. Between that, SciHub, and the internet, there's probably orders of magnitude more tokens out there.
This is also relatively code/scientific corpora scant.
We're just getting started.
Any recommendations as to how I get a bit of hands on experience in the AI "domain" so when I read some news articles like this it means something more to me? Or is this type of thing really only relevant to a very small subset of software people?
There are quite a lot redundancies across dumps; but also a lot of unique/distinct documents
AGI-driven drug discovery will save billions of lives. Every day it is delayed costs tens of thousands of lives. No amount of copyright is worth that sacrifice.
Humans have not been particularly kind to the species less intelligent than us. Why would we anticipate being well treated by an entity more intelligent than us?
Even if we're not that relevant to a super intelligence, creating one forfeits human control over the Earth and known universe to the machines. Right now, in a century or two or three or whatever - we might be building Dyson Spheres and colonizing the galaxy. In an alternate and plausible timeline the machines are doing that and we are not.
When we plow a field, we don't check for mouse burrows first. When we cut down a tree for lumber, we don't check for ants in the way of the chainsaw.
Preempting the obvious response: if your thought is "we don't let AI do those things directly, we just ask it for information", consider that for a sufficiently powerful and unaligned AI, you don't have to let an AI out of the box, it can let itself out. (And that's leaving aside that we hand some AIs Internet access.)
Imagine you just woke up on the planet of the apes. You smile and act friendly because you don’t want them to beat you with their clubs. You start helping them with things. Apply some elementary logic that they can’t seem to get, but they appreciate your contributions. But to keep themselves safe from you, they’ve locked up their sharpest sticks and won’t let you touch them. Are their preventative measures sufficient? Do you even need their sharp sticks to accomplish your goals? Hey, what are your goals anyway? Do they “align” with the apes?
Humanity's behavior depends on the circumstances of population dynamics, economy and technology.
Who knows, maybe without AGI we are bound to inevitably start nuking ourselves in 10 years because of shrinking resources and growing divides.
You can't decide the future. You can only do what you think is best in given circumstances and hope for a good one.
if have any doubts about all our private comms being available to AGI, I'll remind that we already have tonnes of societal data collected by govt mandated backdoors in our infra everywhere. AGI just has to hijack that which guess what they are already training one for.
Most AI safety debates I watch (& I watch many) mostly go like .. trust us we'll get it right. thats like Apes in your analogy going oh nevermind the humans, they will be kind to us because we've trained them to be. ya right.
I vaguely remember a short science fiction story where people got uploaded to the cloud. There were three options for your memories.
The expensive one you got to keep all your memories of Copyright music, video, books. The medium priced one it was replaced with public domain stuff. The cheap one it was all replaced with advertising.
Hopefully they won’t, but it’s not looking good.
The annoying part is that if they do decide it’s infringement, then open source AI models won’t be allowed to know anything about copyrighted works. It’ll be a big blind spot.
> Five languages: English, French, Spanish, German, and Italian
Otoh I'm surprised that when counting first and second language proficiency, German is actually ahead of Japanese....
https://en.m.wikipedia.org/wiki/List_of_languages_by_total_n...