RedPajama: Reproduction of LLaMA with friendly license
together.xyz
together.xyz
(While it's true that the actual model code of Llama is properly open source, it's also useless for inference by itself. Claiming these models are open source seems like having your cake and eating it too - you get accolades for "open sourcing" but still get to control what happens with it.)
While RedPajama has yet to commit to a license (from what I can see, it is late at night…), they are making all the right noises and I am hopeful that my prediction that we are about to see the floodgates of truly open models blow open and that OpenAI’s “moat” will be proving to be a lot shallower than what they and many others have made us believe over the last six months will come true.
this is a required statement to conform with China’s constitution, or the superseding authoritative social contract there.
think of it like if the Patriot Act was an article of the constitution instead of a random law subservient to the constitution, it would negate other parts of the constitution that we hold near and dear.
this is a useful similarity as both constitutions have assurances of free speech
just one has a fatal heavily leveraged clause that undermines all other parts of that constitution and dictates all facets of life
Regardless of the cause though, the clause flies afoul of any definition of open out there.
but its not likely that any uncontrollable LLM can start spitting out accuracy or things unhelpful to Beijing’s ethos there and be allowed to operate
the model or the service filtering the model has to be controlled
But doesn't this mean the model training data also excludes anything critical of China?
For example, does their training data include things like this: https://en.wikipedia.org/wiki/1989_Tiananmen_Square_protests... ?
Either way, huge fan, it would be awesome to have a LLaMA set of weights that are fully open.
We are appreciative to the work done by the growing open-source AI community that made this project possible.
That includes:
Participants in building the RedPajama dataset including […] LAION.
Meta AI — […].
EleutherAI — This project is built on the backs of the great team at EleutherAI — including the source code they provided for training GPT-NeoX.
An award of computer time was provided by the INCITE program. This research also used resources of the Oak Ridge Leadership Computing Facility (OLCF), which is a DOE Office of Science User Facility supported under Contract DE-AC05-00OR22725.”
The answer to your question is right there at the bottom of the page in the linked-to blog post :/While I am not a lawyer and Apache 2.0 is likely to be unproblematic, I always find it puzzling as to why people recently are opting to license non-software using software licenses (Apache 2.0 in particular). Hopefully you have access to sensible lawyers, but I was always under the expectation that model weights would fall under a license such as CC-BY rather than Apache 2.0. Sadly it has been too long since I read the recommendations and justifications for this, so I can not find a good reference, but seem to recall the advice came out of FSF.
I don't think the "Open Source" label people are using is accurate, and I heavily agree that a common thing that companies seem to be trying to do in this space is release what are essentially closed models while calling them open, and it's a really dangerous direction for AI to go. So nothing in your comment is wrong.
But it also feels a little bit like ceding ground to just assume that Llama can't be used commercially just because Facebook says it can't. I never signed a EULA with them, that claim depends entirely on whether or not model weights are under copyright (or under some similar form of IP protection, some people have brought up trade secrets).
And I don't have a super-strong opinion necessarily, but I'm not sure that's a safe assumption for people to make, and I kind of think it might be good to throw an asterisk next to "can't be used for commercial projects" whenever we talk about Llama's restrictions.
But again, I agree with you, it's not the same as saying Llama is Open Source. Even if it does get ruled as having weaker protections, I don't think the term would really apply.
On that subject, to the best of my knowledge, I also haven't signed any kind of agreement with OpenAI. I've done all of my GPT testing through 3rd-party services or portals that don't require signing EULAs to use.
And then coming second appears to be "companies and whoever who seek to make money, and intend to make some sort of legal restriction part of the biz model."
I have no answers or even predictions here except "this is gonna be interesting."
I ran a HEAD request against them all to sum up the total file size, and it's 2.67TB total.
Here's a Datasette Lite URL that lets you explore the size metadata about those files: https://lite.datasette.io/?json=https://gist.github.com/simo...
And a SQL query that shows the breakdown across the different sources:
https://lite.datasette.io/?json=https://gist.github.com/simo...
Sizes here are in GB:
common_crawl 1341.6166818914935
c4 806.7667234372348
github 212.1786002581939
wikipedia 111.89125544670969
book 100.43162744678557
arxiv 87.35323827341199
stackexchange 74.54870238155127
Common Crawl is in there a few times - they have the following folders: common_crawl/2020-05 198 files
common_crawl/2021-04 176 files
common_crawl/2023-06 175 files
common_crawl/2022-05 157 files
common_crawl/2019-30 153 files
And then C4 as well, which is "a colossal, cleaned version of Common Crawl's web crawl corpus. It was based on Common Crawl dataset": https://paperswithcode.com/dataset/c4So ~4x the size of the Pile, any idea how it stacks up in terms of quality to other big datasets?
In your opinion do you think this will hamper the model at all? Or is it still more than enough to get good coding assistance?
I wonder how hard it would be to fine-tune something built on RedPajama on further code examples to improve performance there.
Compute optimal here means the point at which it makes sense to move from a smaller to a larger model assuming that: (a) you have a fixed compute budget of FLOPs, and (b) you want to train the best model possible. The problem is that this applies only to training and assumes nothing about the cost of inference. If you actually need to deploy these trained models and support them long-term for hundreds, thousands, even millions of people to use, would you rather deploy a 13B model or a 30B model at the same level of quality, even if the 13B model would be more costly to train?
There is going to be a point at which these models plateau and further improvement will not be possible without moving to a larger model, but Llama doesn't get there quite yet.
I'm not sure if this is the right place to ask about this, but could you consider training an LLM using a more advanced, sparse transformer architecture (specifically, "Terraformer" from this paper https://arxiv.org/abs/2111.12763 and this codebase https://github.com/google/trax/blob/master/trax/models/resea... by Google Brain and OpenAI)? I understand the pressure to focus on training a straightforward LLaMA replication, but of course you see that it's a legacy dense architecture which limits its inference performance. This new architecture is not just an academic curiosity but is already validated at scale by Google, providing 10x+ inference performance boost on the same hardware.
Frankly, the community's compute budget - for training and for inference - isn't infinite, and neither is the public's interest in models that do not have advantage (at least in convenience) over closed-source ones; and so we should utilize both those resources as efficiently as possible. It could be a big step forward if you trained at least LLaMA-Terraformer-7B and 13B foundation models on the whole dataset.
https://yaofu.notion.site/How-does-GPT-Obtain-its-Ability-Tr...
I think it's essential to increase the quantity of code tokens.
wget -i https://data.together.xyz/redpajama-data-1T/v1.0.0/urls.txt
This is also mentioned on the dataset card for redpajama-data-1T on Huggingface [1].[1]: https://huggingface.co/datasets/togethercomputer/RedPajama-D...
The data looks like it should compress pretty well. If you use something like btrfs's transparent compression, I wouldn't be surprised if it all fit in less than 0.75TB of disk space while still being usable to any tool that expects uncompressed data.
Edit: It looks like some of this data is already compressed, so maybe not.
Doesn't this imply the produced model has to be CC-BY-SA too?
CC-BY-SA content needs attribution too, but I don’t see the(se) model(s) in the current state being able to do so.
I imagine we’re gonna see the IBM PC bios/Unix/ReactOS “tainted code” arguments again in court, this time is not the human who is more-or-less knowingly responsible for sneaking in copyrighted code.
The only theory under which training this sort of model is remotely legal is that doing so is not prohibited by copyright law in the first place. If that theory is correct they don't need a license, and they don't need to abide by any terms of licenses that they were granted without asking.
If that theory is incorrect, they have to comply with the stackoverflow license, but they also have to not use any of the (massive amounts) of unlicensed training data they are using, and comply with the numerous incompatible licenses other sources of training data are licensed under. In other words it's impossible to do this.
I tend to favour the view that in this case it is legal (by way of the de minimis doctrine), but I don't think it's a trivial question.
Distribution is when the issue arises - not consumption and construction of a mental model.
I acknowledge the parallels are imperfect and this all needs to be worked out in court. But it’s possible that at the pace LLMs are developing, by the time courts start addressing these questions we’ll already be questioning whether the distinction between machines and people is as big as we thought.
This will definitely accelerate progress in LLM research, productization and safety. Alpaca, vicuna, gpt4all and others are sporadic repesentations of this that could become a continuous improvement process were the LLM and its license truely open source.
An interesting possible side effect of a GPL-like license is that AIs become unlikely to be trained on private data, the usual moat that big tech wouldn't want/just can't make public if it were to use those GPL-like licensed models.
Pythia 13B is worse than LLaMA-7B and requires double the resources.
https://www.databricks.com/blog/2023/04/12/dolly-first-open-...
[1] https://huggingface.co/OpenAssistant/oasst-sft-4-pythia-12b-...
[2] https://huggingface.co/OpenAssistant/oasst-llama-based-model...
The OpenAssistant effort gets an A+ for open source contributions.
https://youtube.com/watch?v=ddG2fM9i4Kk&t=132
It's easy to miss but after the negative build-up he says: "and... I'm kidding!"
Too bad my original comment is too old to edit.
For some technical reason?
https://www.gesetze-im-internet.de/urhg/__44b.html
https://www.gesetze-im-internet.de/englisch_urhg/englisch_ur...
To push a bit further, there’s something that just feels particularly off about assuming everyone’s content is up for grabs unless the producers do the work to opt out. I think there’s an especially palpable bit of irony looking at it from the EU’s perspective—where cookies must be opt-in, but grabbing all your copyrighted material so companies can do whatever they like with it places the burden on the owner to opt-out. It just feels backward. Perhaps one should have to expressly opt-in to allowing their work to be accessible as training data. At least then there will be a clear signal that the producer of the work can’t later complain, as they willingly granted permission.
While you don't need that to generate an image, it's something SD can actual do extremely well with ControlNet, Textual Inversion, LoRA, img2img and so on.
That's an area where things are going to get interesting in the future, as you can take any image, feed it into SD and produce hundreds of AI images from it. Very easily, without much effort and within minutes. The delineating line between original work and derivative becomes extremely blurry here, as what you are copying is not "the image", but just concepts within the image, that can be a pose, camera angle, scene layout, art style or really anything. You can "copy" it with as much variation as you want, you can remix it with other images, text prompts and so on. Where does "looking at reference" stop and "doing a copyright violation" start?
The spooky part with AI art that it stops images from being singular entities, with AI you can explode every piece into millions of possible variations. AI is so fast at generating art that a future where we could generate movies in real time might not be far away. It's already fast enough to produce images and text stories faster than you can consume them. There might be a fundamental shift in art consumption ahead of us.
"A reservation of use in the case of works which are available online is effective only if it is made in a machine-readable format."
It's also far from clear to me whether a court would find training an LLM to constitute text and data mining 'for the purpose of gathering information, in particular regarding patterns, trends and correlations'.
It's pretty normal for a law to not be specific on the technicalities so they don't have to update the law whenever the software changes. The de facto standard to prevent bots from scraping your sites has been robots.txt for almost 30 years.
If artists didn't mind Google scraping their images, putting them on their site, adding ads and making billions, I really don't see them having much of a justification to call out StableDiffusion for "stealing" their stuff. In general artists would be in a lot of trouble if taking stuff from the Internet would be outlawed, as that's where they get all their reference images from too.
Either way, I am sure we'll see quite a few lawsuits going forward, laws are always open to interpretation, especially when new technology archives. But long term I really see copyright in general being in a lot of trouble, since derivatives and remixes are becoming completely trivial with AI. Where does the original work stop and the copyright violation starts is being rather difficult to decide when you can just wander around latent space and create literally thousands of similar images in minutes, with as much or as little variation as you want.
While a court would likely conclude that a watermark on an image is not 'machine-readable' (I say likely—OCR technology would however make it possible that a court could find that a watermark is machine readable), I would say that because the law does not require a specific method, I think it might be found that a copyright notice in the footer, or in an image caption, is indeed 'machine-readable'.
On balance, I agree that there's a lot of things we are woefully underprepared for coming up in the very near future on using tools in this way to generate art. The answer is not simply to try and lock up all the art away from the robots—but I don't know what the answer actually is.
Also, server/hardware costs are still a limiting factor for running and finetuning the larger 33/65B Llama models. Especially, if they can only be used for personal toy projects.
Midjourney is way more versatile than SD. If you start getting some fine tuned models on civitai, trained to do well some specific tasks, you can get comparable quality but I haven't seen a single model which is able to replace Midjourney.
Llama is no different, it has ok performance on generic queries but still far away from GPT3.5: if you start fine-tuning you can get good perf on specific tasks.
They’re very similar offerings if you’re willing to put in the work on SD.
Sure, its very easy to get good results fast, but the tuning that avoids "uglier" images is the same that removes a lot of versatility compared to SD
Also controlnet is a killer feature
Sure, Midjourney is a centralized commercial service with a clear statement that you (as a paid user) own the images you create. While that doesn’t resolve all potential copyright issues (as there are still at least theoretical issues with the underlying dataset), if you doing something commercial with it like, say, a webcomic from which you derive income, its a lot simpler than dealing with the SD ecosystem where the plethora of models also have different stated usage restrictions, different suppliers (many of which are hobbyists) to keep track of, and more potential avenues of indirect copyright risk, as well. For some webcomics, even the base CreativeML Open RAIL-M license itself might be problematic.
This isn’t a technical or quality advantage, but its definitely an advantage that would very often tip the balance between two tools if both are minimally adequate to your task.
Models on CivitAI are okay. Cool if you're looking for a certain style and/or want to create something that looks like the training images but style isn't everything.
Midjourney generates much better than "slightly better images" and the very fact you say this just tells me you've not even used the thing in any real capacity.
I am the author of submissions such as: https://news.ycombinator.com/item?id=35181433, and I am one of the people responsible for the enthusiasm behind the performance of MJ v5.
But no, MJ is not much better if you know how to use SD, although if what you did with SD was just put a prompt in a huggingface space, I can understand why you say that.
>I never said you can't do impressive things with SD, but feel free to share these comics.
I am arguing that they are better than any comics made with MJ, not that they are simply impressive, that's really the entire point. I know some on Pixiv, you can look them up if you want; I am not linking them for obvious reasons (to say they are NSFW is putting it mildly).
I'm the person behind these - [url-redacted]
I think it's safe to say i know something about SD's capabilities.
>I am arguing that they are better than any comics made with MJ, not that they are simply impressive, that's really the entire point.
Sure that's why i'm asking you to link these comics that are supposedly better than anything Midjourney has ever produced. With a claim like that, i'm sure you understand wanting to see results.
>You can go look them up on Pixiv if you want, they host some; I am not linking them for obvious reasons (to say they are NSFW is putting it mildly).
So you can't link anything that isn't NSFW on pixiv? Lol, that just solidifies my point. Frankly if the best you can come up with is pseudo porn(or maybe not pseudo lol) on pixiv (i don't imagine any readers of that will care about the things i'm looking for) then that's not a very good look.
https://www.pixiv.net/en/artworks/107271972
Though, if they're training off official character art that's less cool than reinterpreting it themselves. Means you don't have a "house style".
I think you're being a bit generous there. Either I'm using it seriously wrong or SD can only generate vague blobs while Midjourney can make some proper stuff. It's a larger difference than GPT 3.5 vs GPT 4.
You are definitely using it wrong, if the alternative is “SD can only generate vague blobs”. Even the base SD models are much better than that (though, the strength of the SD ecosystem is the availability of custom checkpoints, hypernetworks, LORAs, embdeddings, ControlNet, etc., not just the base models.)
Not sure why there's even an option to go below 512.
1. It's a large corpus of technical knowledge; 2. The language is written by experts in a field and reviewed many times, and 3. They have technical drawings with labels and references in the text
The only downside I suppose is that sometimes patents are written with "just enough knowledge" to get it granted but not too much to give away the secret sauce. That's not really that different from many scholarly papers though.
To give a size of scale, the granted patent texts of 2020 (without images) is about 160 GB of data, and we have digitized grants going back to at least 1970.
In fact I'm downloading a whole batch of patent texts right now because I wanted to experiment with semantic search on patent texts.
Anyone here have any pointers on what the state of the art method for semantic search through a large corpus would be? I've just started researching and BERT and friends seems like it was popular about 2 years ago but things move so fast I wouldn't know what I should do now.
What about a medium sized corpus of text, say 100.000 pages of text?
Part of its contents come from the "USPTO Backgrounds" dataset. From The Pile's paper:
> USPTO Backgrounds is a dataset of background sections from patents granted by the United States Patent and Trademark Office, derived from its published bulk archives. A typical patent background lays out the general context of the invention, gives an overview of the technical field, and sets up the framing of the problem space. We included USPTO Backgrounds because it contains a large volume of technical writing on applied subjects, aimed at a non-technical audience.
More details in the paper: https://arxiv.org/pdf/2101.00027.pdf
The Pile: https://pile.eleuther.ai/
Some links: https://github.com/bovlb/opencyc https://github.com/asanchez75/opencyc
That said, it certainly seems like there hasn't been recent work on hosting the OpenCyc knowledge graph in a reasonably modern way, much less the more recent closed-source work by Cycorp (https://cyc.com/). And it's likely GPT-4 doesn't know the full capabilities beyond whatever tutorials were on the web at the time of its training. If I were Cycorp I'd be seriously looking at developing this kind of hybrid model, with an agent model having access to recall their closed-source examples, as a paid cloud offering; there would likely be many who would desire this best-of-both-worlds.
Given the immense momentum behind LLaMA, I'm pretty disappointed that Meta won't just open-source it, but I guess reproducing it is better long-term.
I remember setting up my PS3 & home desktop for folding project. Its fair game especially if I can use the box to heat the room instead of the furnance.
Surely there is a better alternative than a bunch of A100s on AWS...
Here is a much more creative reading by Ludacris [0]
An order of magnitude lower GPU-hour time, plus if you train it for 210 days instead of 21 days, means you could do a 7B model with 20 consumer GPUs which are $1000 apiece. $20k, not counting mainboard, etc. Really not bad. Might even be doable as a volunteer project.
Also most training is done using bfloat, not single precision (which is usually only used for accumulators)
When training a 65B-parameter model, our code processes around 380 tokens/sec/GPU on 2048 A100 GPU with 80GB of RAM. This means that training over our dataset containing 1.4T tokens takes approximately 21 days.
At $4/GPU-hour per A100 80GB GPU, that's $4 * 2,048 * 21 * 24 = $4,128,768.
> The objective of the scaling laws from Hoffmann et al. (2022) is to determine how to best scale the dataset and model sizes for a particular training compute budget. However, this objective disregards the inference budget, which becomes critical when serving a language model at scale. In this context, given a target level of performance, the preferred model is not the fastest to train but the fastest at inference, and although it may be cheaper to train a large model to reach a certain level of performance, a smaller one trained longer will ultimately be cheaper at inference.
Obviously this would increase training costs substantially, so I understand why more languages are not included in the base dataset.
Sounds like they already have the compute and began training.
1. Getting data of equal or better quality
2. Securing the funding/hardware required for training
3. Learning/figuring out the training challenges needed to tune the process (the PhD part)
It seems #1 is the relatively lowest hanging fruit and a prerequisite for the other two, and that's what the project is (rightfully) tackling at this stage. #2 could be solved by many ways, and doesn't require much innovation if the project and the team are solid. Which takes me to #3, which on the other hand seems to be the make or break part of the project.
I'm not one to doubt the technical prowesses of the RedPajama's team and their contributors, I rather see it economically. How can an AI open-source project compete with big tech in attracting the brilliant minds of our generation? It's enough to look at levels.xyz to see the battle is not ... level.
There's a serious economical challenge in here to have any sort of sustainable open source initiative in AI.
Therefore, anyone will be able to fine-tune the RedPajama models using Vicuna or other datasets, given they will be fully open-source.
The RedPajama instruction-tuned models will be fine-tuned only with instruction labels from human labelers and OpenChatKit feedback (). We feel this will keep these models fully "clean" for use in commercial applications without using the output of other commercial models like were used in Alpaca or Vicuna. However, we'll be excited to see all the great fine-tunes created by the open community and are eager to see how close open-source models can get to the quality of leading commercial models over time!!
() OpenChatKit: https://huggingface.co/spaces/togethercomputer/OpenChatKit
I wonder... who is paying? Will there be restrictions like ethics clauses and suchlike. Not necessarily a bad thing if they do. Will there be restrictions on commercial use.
Is 40 a100s enough though? I am interested in what this would cost.
> When training a 65B-parameter model, our code processes around 380 tokens/sec/GPU on 2048 A100 GPU with 80GB of RAM.[1]
Note that you probably need to budget for double to triple that because things go wrong and it usually takes multiple starts to get a good training run.
Smaller models are cheaper though.
2.5G filtered_08cdfa755e6d4d89b673d5bd1acee5f6.sampled.jsonl
834M filtered_08cdfa755e6d4d89b673d5bd1acee5f6.sampled.jsonl.lz4Then fine tuning came from training on actual python code on GitHub.
At the model understands the python documentation and the implementation standard library/interpreter. Then is there a reduction of data needed for code generation therefore reducing the size of the data set used for code generation?
I do wonder if anyone is considering mixing in larger and larger percentages of The Stack https://huggingface.co/datasets/bigcode/the-stack with this or the Pile to get more code and see what happens.
(Likely beyond mere mortals' budgets though.)
Does anyone know which licenses are filtered into the dataset?
[1] https://huggingface.co/datasets/togethercomputer/RedPajama-D...
In self-supervised learning, the training target is a modified version of the input.
Transformers are trained in a self supervised manner. The problem they solve is "given a sequence of N tokens, what is the next most likely token?"
No labels required :)
Models like gpt3 get turned into models like ChatGPT through RLHF (reinforcement learning from human feedback), by fine-tuning the model further on prompts in the style we'd like them to respond in, typically
User: question
Bot: Response
This is done by handcrafting or modifying data from places like stack exchange.