The Pile is a 825 GiB diverse, open-source language modelling data set (2020)
pile.eleuther.ai
pile.eleuther.ai
"So here’s the big picture. There are three sets of datasets: 1. Data exists out there in the world. It has been collected into datasets and posted online. I’ll call this raw data. 2. We take that data, clean it, and process it for language modeling. I’ll call this per-set data. 3. We combine those per-set data into one massive dataset, the Pile. This is heavily processed, including weighing the components.
We created 2 and 3 and put them online. We put 2 online so that people can reweigh and remix the data if they wish, but we expect most people to just download 3 and use it out of the box. Access to 3 will be provided in several forms, including HuggingFace and from our website.
2 and 3 are not copyright violations, even if the data is copyrighted, because they fall under fair use (at least in the US).
The Pile contains code that turns 1 into 2 and code that turns 2 into 3.
When you download Maroon 5 from a website, you are creating a dataset corresponding to 2. That can be copyright violation depending on what you do with it, but our use is not a copyright violation."
In terms of fair use, one of the larger factors is the 'market substitution' factor, which basically means "does this use compete with otherwise licensed uses that people would ordinarily pay for?" AI absolutely does compete with human artists for the same market. In fact, it's winning handily[1], because you don't have to pay human artists. AI art models absolutely shouldn't be trained on anything with copyright on it.
The other factors don't fare much better. Nature of the original work will differ based on the plaintiff, but the purpose and character of the AI's use of that work is very much commercial. And the amount and substantiality of the use is complete and total. I don't see AI being fair use - at least, not in every one of the many, many training lawsuits currently ongoing against OpenAI and Stability.
[0] Starting with any body of text, an LLM, and an empty context window, compute the next-token probabilities and take the highest one. If it matches the source text, output a 1 bit. If it doesn't, output 0 followed by the ID of the correct next token. Add the correct token to the context window and repeat until the text has been fully compressed. This produces a list of perplexities (wrong words) for the given text which can be used to guide the LLM to output the original work.
[1] Hey, remember when both WotC (biggest art commissioner on the planet) and Wacom (hardware vendor that sells art tools and payment terminals[2]) both got caught using AI art after making very loud and public pledges to not do that? They both wound up buying stock photography on marketplaces that are absolutely flooded with AI trash.
[2] All the credit card readers in Japan are built by Wacom, which is really funny as an artist
And that's legitimate inventions I'm talking about. Just wait until the patent trolls figure out how to get the A.I.'s to divulge patent violations to sue the A.I. suppliers.
From the ruling:
> Assuming the truth of Plaintiffs’ allegations - that Defendants used Plaintiffs’ copyrighted works to train their language models for commercial profit - the Court concludes that Defendants’ conduct may constitute an unfair practice.6 Therefore, this portion of the UCL claim may proceed.
https://caselaw.findlaw.com/court/us-dis-crt-n-d-cal/1158180...
So in that sense I understand the response that "they don't violate copyright" by studying the material. Again, I don't pretend to be a lawyer, and not every law has to follow my logic.
This isn't about the output for content generators or about the abstract numeric weights that they operate over. That's more complex and a largely open question.
But this is literally about indiscriminately distributing copyrighted works in a large, convenient archive while arguing that it's okay because you normalized the formatting a bit and because you suspect that some people might find "fair use" value in it.
Meanwhile the complainant (such as GRRM) forgets that often passages from said articles and books are strewn throughout the Internet. Of course chat GPT can drop passages from the GoT books; there's several entire fucking wikis for that franchise that reference passages, quotes, details etc.
Same goes for news articles, passages of which are often quoted by other sources or websites.
Not that chatGPT has reproduced many works in whole, but it's an interesting logic problem for fair use law: if I have copyrighted article X, but websites ABCDEF all quote various passages of my article (ie fair use, critique, etc) and then ABCDEF is used to train an LLM, if the LLM can _reassemble_ the article from quoted passages without referencing article X itself, is it copyright infringement or fair use?
In the case of The Pile, "processing for language modelling" means "converting epub and pdf into plain text, maybe deduplicating, maybe removing some sorts of detectably malformed files"
So not a particularly lossy conversion.
Then "they're doing something amazing, they don't need permission, and the cat is already out of the bag, and similar musings".
Seriously, it's both copyright infringement, and unethical. This is why I don't use any of the popular AI tools, or even AI add-ons in Evernote, Notion, etc. They all link back to the usual suspects.
Open science repositories would take down the "dataset" immediately (or at least limit its access) if a copyright holder brings the matter to the eyes of the admins.
Ah, the Uber theory of law. Works surprisingly well for some reason.
I can see the problem where direct and faithful replication is possible but where it isn't is there still a problem? Or is the automatable aspect, the scale at which it can occur, that is the problem?
An AI system consumes something perfectly, then ingrains it into its weights perfectly, and becomes capable of imitating the same thing perfectly. Plus, ther are no other internal or external factors which affect these "generation" over time. Hence, it mixes and reproduces based on what it consumed, solely.
I might get inspired by people, and add my own values to it, iterate over it and diverge from what I'm inspired from ultimately to create my own style. AI doesn't work like that. Also, if I do the same amount of inspiration with the same precision and accuracy, I'll be neck deep in accusations and lawsuits (for the right reasons).
As a result, just because we fail to ask the right questions to reproduce the training data verbatim or almost verbatim doesn't mean that the information is not there. At the end, a neural network is a compression algorithm which encodes data in terms of weights. Given the correct input, you can regenerate the training data as is.
Unless you have special abilities, you can't read 500 books an hour, remember them perfectly, and generate derivative works by mashing all of them together. If I do and try to sell a novel, I'll be ridiculed no end. If I write a Ph.D. the same way and try to defend it, I'll be banned from academia for three lifetimes at least.
For more elaboration on the subject, see [0].
If you divide Stable Diffusion's file size by the number of images used to train it, you get something like 1.2 bits per image, and it is physically impossible to get this kind of a compression ratio.
The actual problem with AI is that it sometimes plagiarizes random fragments of the work it is trained on, even if it is not the user's intend, and we currently don't really know how to fully prevent this.
Same for code generating models trained on Open Source and Free Software. Tons of licenses violated, from strong copyleft to source available models and reproduced (almost) verbatim with comments intact.
Some researcher's codebase is almost completely reproducible without any licensing information just by hinting the function names.
Maybe for the image compression it's borderline impossible for now due to network size, but for text and code, generation of training data almost verbatim is very possible and straightforward.
Also in image generation models, style transfer is the bigger problem, because it completely eliminates the artist who uses/created the style in the first place. "You pioneered this, and we fine tuned this model with your images, and we can do your work for free, without you, have a nice day". However, the artist's life expenses doesn't disappear when they're transferred to an image generation model.
This is also unethical.
Just because you can doesn’t mean you should. That's what I'm trying to say.
IIRC it was like 1.4 bytes before adding in random initial and prompt. And Amiga Four-byte Burger is 4 bytes long.
[0]: https://bytecellar.com/2023/04/24/lost-amiga-four-byte-burge...
call me skeptical, seeding a torrent of movies that you downloaded from elsewhere on the internet isn’t “fair use” and the pile isn’t just code for transforming data, it is the redistributed data itself
by this logic i could legally run a libgen mirror
For those concerned, I have an article that covers legalities, each dataset (including The Pile), legal issues with them, alternatives that are legal, and a copyright amendment that balances all sides.
http://gethisword.com/tech/exploringai/
Looking back at my proposal, I think we need at least three rules passed immediately in at least one country:
1. All copyrighted works can, if a person has legal access, be used for training AI systems. Any terms restricting copyrighted works from use in training, charging more for that, restricting downloads for it, etc are illegal. Every act of publishing can benefit both a human mind and AI training equally.
2. People can copy and transform for their own use any work they have access to only for AI training. This might include reverse engineering for extraction, multiple copies in different formats, and so on. They can do whatever is needed to get it into the AI system. Other uses or abuse of this data is subject to existing law.
3. Any work published online for free and with public access can be copied, shared, processed, and bundled for AI training. That’s regardless of its terms.
Note: In No. 2 and No. 3, the resulting AI’s copyright will be determined by existing law about AI’s and mixing copyrighted works. Or no copyright if that’s the law.
4. If AI outputs are copywritten, their status will be the same as if the user published it themselves while relying on prior works. AI training sets will also be public to determine this.
With those rules, we can share works like those in The Pile, still pay creators that want to be paid, be less likely to just steal existing work, and infringement in outputs is still illegal. What do you all think of that?
This cannot be known until it is litigated. Fair Use is not something you can unilaterally declare and have it be so, just like you can't be like Michael Scott in the Office shouting "I declare bankruptcy!" OpenAI is currently defending itself against the New York Times for this very reason.
There's a multi-factor test that courts weigh the facts against in making a determination as to whether a prima facie copyright violation would be protected under a Fair Use defense:
Factor 1: The Purpose and Character of the Use
Factor 2: The Nature of the Copyrighted Work
Factor 3: The Amount or Substantiality of the Portion Used
Factor 4: The Effect of the Use on the Potential Market for or Value of the Work
See https://copyright.columbia.edu/basics/fair-use.html for a pretty good overview of what the analysis entails.
By no means am I an expert in copyright law, but factor 3 seems like very bad news if you're OpenAI.
When a court determines that it isn't, you can continue to argue it as much as you like (to deaf ears), and yet you're still liable to the copyright holder. Whether it's "incorrect" or not is then irrelevant. Let's not argue semantics here.
No, they are different. "fair use" is a legal term. It is not like saying "I use it like this, I think it is fair!", the term "fair use" literally is a legal term that means a particular thing in the court of law.
Correct, but what isn't clear here is their rationale for why they think they're covered by fair use. Does anybody have that information?
I'm not saying their interpretation is correct, but seems to be germane to this discussion. The parent comment seems to assume none of this has been litigated yet, which might also be true. Or not.
I'm open to the argument that generators built with models that consumed copyrighted data may evade copyright obligations on their output, but surely the data sets themselves are bound by any copyright on their content?
Sorry that's not how damages are calculated in the US tort system.
I also think that the case is different here, since in your example, there is a specific amount of money being stolen, while in the books3 case, there is an unspecified amount of money not being made by the authors.
1) pay the author,
2) implement guaranteed citation of the author any time the model gave an answer that was directly derivative, with an option to not do so if the summary was sufficiently vague, or
3) ignore the author's book completely as training data
we would all choose 3).
Which is all to say that information shouldn't be hoarded and guarded. If it can produce something more than the sum of its parts we should use it to do so. The result of that should, on the same grounds, not be hoarded and guarded, doubly so being based on the work of others.
The solution to wealth disparity cannot include "invent menial untalented high-paying labor for people to do".
There's no difference between a machine replacing a human writing a book and a machine replacing a human making a piece of wooden furniture. Literature and craftwork are equally rewarding and beneficial imo.
And people still make wooden furniture by hand even after IKEA. If someone wants cheap and uninteresting they'll buy IKEA, if someone wants an interesting and unique piece, or to support handmade things then they'll buy from a woodworker.
Thank you. You know in some ways it's an honour and a privilege to live in such times of progress. The very act of publishing is to "let go", and hope that your words and ideas contribute to something bigger and beyond your life. I never believed much in "intellectual property" as it's all stuff that flows through us.
> I hope that the fruits of their labours are returned to them
They rarely are, because knowledge and creativity are not greatly valued in our time. But authors, artists and scientists go into that with eyes wide open these days. The rewards come in other ways, as the more you give and put into life the more you get out.
> rather than being selfishly hoarded by the few with the resources necessary to produce those fruits
This is not what we fear. Hoard away. We will simply take back what is ours, whenever we desire it. The hoarders will never win against what they call "piracy", because they have no moral right. In the long run, they are on the wrong side of history.
Far worse, and more likely is that the creative and technical works of generations of artists and scientists are going to be turned to exactly the opposite of what they would want. They will be used to harm and disempower humans, divide society instead of heal it, and even make the pursuits of art, science and knowledge irrelevant.
We cannot take back our words, or our formulas, or our paintings or our songs. But we can take back tech.
It might be something that's part of their job, requires citations since they value credit, optionally bans commercial use, and maybe has a patent. Arxiv papers are a mix of that. Many on StackOverflow and Hacker News want attribution with some asserting copyright in their comments. Film producers usually want the subtitles to accompany sales of their movies. For FreeLaw, the material is public by law about people who might have never even wanted to testify or be remembered. FreeLaw itself was seeking donations for its service with an additional request to protect the names of people in the dataset. StackOverflow's license explicitly bans copying their data without permission with many individual users also wanting self-promotion to happen side by side with their answers.
So, many things that are in The Pile are works people published hoping to gain some benefit in return. They often had terms that banned their reproduction in ways that prevented them from getting that benefit. They were just public for humans to read and learn from. Some allowed sharing but just wanted credit.
The A.I. users of these works ignore all of that by taking what they made conditionally available without meeting the conditions that benefit the authors. Whereas, if you got the same content on Hacker News or Arxiv, the authors might benefit from it. Even a pirate would benefit them more than A.I. companies because the users would often at least know the author or source site. So, the fruits of their labor were taken, not given, by those who are the least beneficial to them.
I will note that some people do publish truly free content that has no strings attached. Mostly public domain or CC-0. Those are exceptions that might fit your description.
Throw a dart at a wall filled with every notable author/publisher ever and whoever you hit probably owns some of this data.
Apparently you can just do whatever as long as you say it's for AI research, go post Blu-ray rips online, it's fine provided you have a .ai domain :^)
copyrights do expire, and any books older than Mickey Mouse are public domain, so it's not every notable author ever
Bram Stokers bones will be relieved to hear that their work isn't being misappropriated.
If you meant transcribing dialogue from a TV show is violating copyright, I'm not so sure, it's relatively common to quote dialogue for varied purposes, ex. TV critics
Definitely understand if you're saying the whole dialogue for a TV show is copyrighted, but I'm curious about the opensubtitles part, used to work in that area.
n.b. not being argumentative, please don't read it that way, I apologize if it comes off that way:
Not every derived work is a copyright violation, that's why subs and dubs don't get kicked around, you can quote dialogue in an article, etc.[^1]
Answering if it applies to AI is playing out in court currently with ex. NYT v. OpenAI[^2] and Sarah Silverman et al v. OpenAI[^3] and v. Meta.[^4]
[^1] "Copyright doesn't protect against all use of the work or use of derivative works. There are a few exceptions that fall under what's commonly known as the fair use doctrine:" (https://www.legalzoom.com/articles/what-are-derivative-works...)
[^2] https://www.nytimes.com/2023/12/27/business/media/new-york-t...
[^3] https://www.theverge.com/2024/2/13/24072131/sarah-silverman-...
[^4] https://www.hollywoodreporter.com/business/business-news/sar...
There... is a torrent.
Anyways...
Is RedPajama 30T and The Pile "all you need" ? ;)
FTFY.
I can see value, but no labor.
(The authors in best seller's lists may well be immune for a bit longer than others writers, as they're necessarily the top 0.1% of writers, but not forever: nay-sayers claimed that AI could never beat humans at chess or go because the games required special human insight).
The top best selling. Only one of many possible reasons for that might be the quality.
Your argument assumes no marketing is manipulating the best selling lists, something known to happen in new york best selling and others.
Regarding people pretending to read "fancy books" (my term) I think most people just dont read, but I find it annoying, for ex. when people see my shelves, that some people think that I buy those books to impress somebody. It is as if people that do not enjoy reading cannot conceive somebody enjoying it. I think it is slightly antiintellectual, and cynical. I have a better opinion of people in general.
I disagree on two fronts. First, given I responded to "> because authors prefer to be paid for their labor", I think the economics are the key, rather than the artistic merits. The authors and artists suffer economically purely on the basis of the AI doing their work for less, not on the basis of actual artistic merit.
Second, the subjectivity of artistic merit means different people like different things. From a market perspective, this is why horror films get made even though they disgust people like me, it's why kids films get made even though adults outnumber kids, and it's why RomComs exist despite the stereotype of men cringing at them.
You are however correct that it assumes no marketing exists to manipulate the best seller lists. But the marketing is also being outsourced to AI, and I suspect there were more writers writing copy than writing novels and screenplays. Now? Now I'm not so sure, though I'd guess it's still true.
I can sympathise with you about other people thinking you're just virtue signalling with your book collection. What I meant more along the lines of how War and Peace has a reputation for being a book that people like to claim to have read but actually have not, or how many loud atheists state that only atheists have read the bible and that's how they ended up being atheists, though in either case I don't know how accurate the reputations are.
Nay-sayers who said that no possible algorithm could beat humans at chess and go? They were wrong. Nay-sayers who say that these algorithms cannot write better books than humans? Well…
Personally, I'd be referring to the family of algorithms that purely take as input a context window and provide as output a prediction of the next token likelihood. (Plus or minus iteration, to generate strings of text.) Pejoratively, one might call these "fancy Markov chains", though as with most pejoratives, that's overly reductive.
All the approaches we're seeing marketed heavily are just fancy Markov chains. I expect every "new algorithm" for the next 5 years at least to be a fancy Markov chain, because that's what I expect to get funding. (I do expect that some people will be working on other approaches, but only for amateurish reasons.)
You can make anything sound unimpressive if you describe it sufficiently poorly.
And: So many different variations are published every month. There are a good number of people in serious research trying approaches that don't use cross entropy loss (ie, strictly next-token prediction).
I don't know what the trajectory of the technology is over the next ten years, but I am positive no one else does either and anyone who thinks they do is wrong.
Now I'm wondering if, with modern knowledge, you could build a 0.5 MeV heavy ion accelerator with only the things available to a medieval alchemist.
I'm thinking probably yes? Triboelectics can get the right voltage. But how good does the vacuum need to be?
> Nay-sayers who say that these algorithms cannot write better books than humans?
They may be right or wrong in the specific, but I think they're asking the wrong question, too specific.
But: AI was seen as "decades" away from beating humans at go, even 6 months before it did.
I don't know how far we are from them writing award winning novels (awards we care about, it doesn't count if it's an award for best AI), though my gut feeling is we need another breakthrough as significant as the transformer model… but even then, that's only a 1σ feeling.
A blanket removal of copyrighted data would make a bot sterile, boring, unrelatable, and ignorant of culture and common memes. We have amazing AI technology. Let's lean into it and see where it goes.
;)
Haha... just kidding... unless.. ?
To get something interesting you would have to generate an instruct dataset from it. It would have to cover a diverse range of tasks. The completions themselves do not make LLMs manifest knowledge and reasoning, a large and diverse instruct dataset does.
"Books3 is a dataset of books derived from a copy of the contents of the Bibliotik private tracker made available by Shawn Presser (Presser, 2020). Bibliotik consists of a mix of fiction and nonfiction books and is almost an order of magnitude larger than our next largest book dataset (BookCorpus2). We included Bibliotik because books are invaluable for long-range context modeling research and coherent storytelling"
“They’re not books, man, they’re a dataset!”
(/s)
Open source is generous sharing, and ethical. Start to nibble away at those ideals and at what point do you slip into unethical? IMDB was a crowd-sourced database put together by a wide community pitching in small efforts which one guy was maintaining like an FAQ. Then the guy maintaining it said, "It's worth money, I own it, screw all of you." How would people react if this happened to wikipedia? But wikipedia is safe because it's a non-profit... you know, like OpenAI, right?
Fair to say that whether or not this is correct is pretty important to all the outstanding court cases on this matter.
But you still have to get your hands on the copyrighted data legally. It might be legal to scan every book an institution owns, and train off it, so long as those scans are not distributed. But it is probably not legal to scrape copyrighted content off torrents - creating the copy to train with is infringing, even if the model's final product maybe isn't.
[1] https://fairuse.stanford.edu/overview/fair-use/four-factors/
books.google.com has been allowed to copy all the books they can lay their hands on, so long as they don't regurgitate them in full, so it's not really the taking, but any subsequent reproductions. And the effect on the market is insubstantial if the alternative wasn't going to be the equivalent sales.
If you take every paragraph in the Harry Potter saga and sort the paragraphs in alphabetical order, it's just as good for training short-context-window models, but not a "harm to the market" leading to a lost sale for anyone who wants to read the books.
I'm pretty sure that's what the lawyers will argue in the Silverman case for example. It's going to be interesting to see how the courts decide.
In the present moment on the other hand, it is the entities in the AI industry (e.g. MS) that have the money and can hire the lawyers to convince the judges. Realistically speaking, it's very likely that things will swing the way of AI companies, which will benefit, albeit indirectly, these guys, even though by themselves they're too small to push their agenda, they're just bit players.
https://the-eye.eu/public/Books/ThoseBooks/Puzzles.tar -- 20-Jan-2023 14:54 -- 6M
and it pretends to be a jigsaw puzzle, but is actually eISBN 9781594868573 - The South Beach diet cookbook / Arthur Agatston
There are ways of copyrighting data, and establishing ownership of intellectual property. Your tumblr fanfic, youtube comments, or HN discussions are not legitimate copyright avenues. Stuff you post to legally scrapeable websites are fair game for fair use.
I can do anything I want in private to any data I collect. I could create an awesome HN LLM on the scraped datasets, and use it privately to my hearts content. I can even set up an API to that LLM that generates content, and, given recent rulings, even if i had all the written copyrighted data in the world, as long as I was making good faith efforts to ensure copyright was being respected and works weren't being recreated verbatim, then I could even use that model commercially. I just couldn't sell it to other people, or distribute it, without entering a different legal regime.
I can collect any data I want from public facing websites.
That's how the internet works; it's how it was designed. There are authentication mechanisms, network configurations, and a myriad other access control schemes you can implement to prevent public access. If you post to sites without those mechanisms, you're tacitly agreeing to give up any plausible claims of protection against a wide array of fair uses well established by precedent cases at this point. If you don't prevent public access, and you've got a domain name on a server, you're tacitly inviting the world to come download whatever it is you have on your server. This is a social good. This is what we want when we participate in the internet.
Insisting on some sort of vague entitlement as to how "your" data gets used completely bypasses the fact that anything you consider to be misused in OpenWebText2 fundamentally stems from the fact that you posted the content to a publicly visible website and gave up any say in what happens thereafter. It was scraped fair and square.
Don't complain that you didn't know the rules, or that life isn't fair.
It's not even clear that terms of service or those little popups on public websites have any legal relevance. If your website is open to the public, then it's fair game. If you post content to a public website, then that content's fair game.
Finally, here's a fun experiment: decide that terms of service don't matter and start building a product by scrapping Facebook or Google. See how they'd react. Actually, no need for guesswork - they clutched their pearls and threatened legal action more than once before. It's a bit of a "have your cake and eat it too" kind of a deal. Their data is precious intellectual property; your stuff is, well, up for grabs.
At any rate - there are ways of staking legitimate claim to content you publish online. Even by doing so, it may not be relevant. Robots.txt is a convention, not a regulation or law. It's respected out of social nicety, not because it's strictly legally required.
If you publish your data to a website where it's publicly visible, you are inviting the world to come download your data. When that data leaves your server and goes to live on the downloader's computer, the downloader can do whatever they want with that data.
It's not clear that it's legally possible to prevent the use of data in training models unless you require someone to sign a contract to that effect before being allowed to download your data.
That would be obnoxious, and I wouldn't bother with your content anymore. Like Instagram, LinkedIn, and Twitter, your site would get a 127.0.0.0 hosts file entry.
The US needs a clear, modern update to copyright law that upholds and maximizes individual rights, as well as privacy and property concerns. We shouldn't be playing this game where we pretend a website is somehow an analogy for a page of text scribed with a quill pen and using laws developed to handle issues when quill and parchment were relevant.
Let's write some new laws where we regulate what things are, and not play tortuous mental gymnastics to contort and butcher existing laws and precedents to say whatever the most expensive lawyers want.
Maybe the social contract allows for people to prevent their conversations from being scraped and used by third parties without explicit consent, even if the conversation is entirely public. I don't like that view, but I see the argument for it.
As things stand, though, fair use and public access make things pretty bright and clear, and rulings in various AI cases so far have favored broad fair use interpretations, and are requiring complainants to show specific, particular harms. If/When those harms are shown, then we'll see if any carveouts will be made, or if broad fair use interpretations will be the baseline for content scraping going forward.
Which, incidentally, the New York Times does and they seem to think they have some legal right to the redistribution of their work.
Maybe they're right, maybe they're wrong, it's up to the courts to decide.
Do be aware that it does include copyrighted content so distribution is piracy.
Also, most image-text dataset pairs contain far worse than that. You might want to check out LAION-5B and what stanford researchers have found in there. Technically, anyone who even touched that could in theory be in some serious, serious trouble. I find it quite remarkable that nothing has happened yet.
- OpenAI and others will just settle with MPAA, RIAA and the likes for a revenue stream (a single digit billion a year, likely) + some kind of control over what people can and cannot do with the AI + the access to the technology to produce their own content.
- artists will see peanuts from the deal, and the big names are going to be able to stop doing any kind of business with artists which are just expenses in their eyes. They will have been replaced by machines that where trained using their art with no compensation whatsoever.
IP is already predatory capitalism, AI will definitely be weaponized against the workers by the owners of the means of “production”.
It's impossible to create such a list while evading all such material.
These hashes is exactly how researchers later discovered this content, so it’s clearly not hard.
Full paper: https://stacks.stanford.edu/file/druid:kh752sm9123/ml_traini...
It is a pretty hard problem.
And this isn’t just “one engineer”. Companies like StabilityAI, Google, etc have used LAION datasets. If you built a dataset you should expend some resources on automated filtering. Don’t include explicit imagery as an intentional choice if you can’t do basic filtering.
That's an amplification of copyright, original expression is protected, but not the ideas themselves, those are free. And don't forget when we actually get to use these models we feed them questions, data, we give corrections - so they are not simply replicating the training set, they learn and do new things with new inputs.
In fact if you think deeply about it, it is silly to accuse AI of copyright violation. Copying the actual book or article is much much faster and cheaper, and exact. Why would I pay a LLM provider to generate it for me from the title and starting phrase? If I already have part of the article, do I still need to generate it with AI? it's silly. LLM regurgitation are basically attacks with special key, entrapments. They don't happen in normal use.
While I don't think it's because you're wrong, per se, it's just that none of this drama really matters.
Somehow people are just not able to get this through their heads. Stable diffusion is like 12GB or something and you have people convinced it's a tool that is cutting and pasting copyrighted works from an enormous image archive.
Good to know I can avoid copyright on a book just by zipping it up!
LLM's are not compressing petabytes of information down to a few gigabytes.
magnet:?xt=urn:btih:0d366035664fdf51cfbe9f733953ba325776e667&dn=EleutherAI_ThePile_v1
I read about it in "The Making of the Atomic Bomb" (1986), but presumably it's featured in the recent movie.
The movie... is a bunch of anecdotes strung together to make a ham-handed point at the end. It was a decent movie if you treat it as a fictional story instead of an actual retelling.
I'd stick with the book. (And if you specifically care about Fermi, I recommend "The Last Man Who Knew Everything" by David Schwartz)
In related news, v2 of the "stack" dataset was recently released
> 3.28B unique files belonging to 104.2M github repositories were collected by traversing the Software Heritage 2023-09-06 graph dataset. Additional repository-level metadata was collected from GitHub Archive data up to 2023-09-14. The total uncompressed size of all files is 67.53TB. Near-deduplication was implemented in the pre-processing pipeline on top of exact deduplication.
V1 vs V2 by Deduped Size Tokens
V1: 2.9TB and 200B
V2: 32.1TB and 900B
I imagine we'll see some fairly powerful open coding models soon. The ones I'm looking at testing are:
dolphincoder-starcoder2-15b-iMat.GGUF
CodeFuse-DeepSeek-33B-iMat.GGUF
OpenCodeInterpreter-DS-33B-iMat.GGUF
starcoder2-15b-instruct-iMat.GGUF
more info
dataset https://huggingface.co/datasets/bigcode/the-stack-v2
gguf quants https://huggingface.co/dranger003
The Stack v2 is ten times larger than its predecessor, yielding a raw dataset of 67.5 TB. Through extensive cleaning, filtering, and subsampling of the source code, along with the incorporation of other high-quality code-related datasets, we created a training set of approximately 3TB (900B+ tokens).
Source: the paper, Section 10 (https://arxiv.org/pdf/2402.19173.pdf)If authors and artists were to join a data union they could do the same thing as studios. If copyrighted law has any real teeth then the data union can send legal requests to whoever is hosting the content and requesting it to be taken down.
I'm not a lawyer but I know the studios definitely do this.
404 Not Found nginx
825 GB is a great candidate for torrent use, whatever was under that broken link better be a torrent magnet.
The Pile: An 800GB dataset of diverse text for language modeling (2020) - https://news.ycombinator.com/item?id=36685115 - July 2023 (70 comments)
Megacorps assume authors, painters, etc are poor and powerless (which lets face it, they are)
we can b** and moan on HN, but megacorps will find ways to use copyrighted works for free.
magnet:?xt=urn:btih:0d366035664fdf51cfbe9f733953ba325776e667&dn=EleutherAI_ThePile_v1&tr=http%3A%2F%2Ftracker.opentrackr.org%3A1337%2Fannounce&tr=udp%3A%2F%2F9.rarbg.to%3A2710%2Fannounce&tr=udp%3A%2F%2F9.rarbg.me%3A2710%2Fannounce&tr=udp%3A%2F%2F3rt.tace.ru%3A60889%2Fannounce&tr=http%3A%2F%2F5rt.tace.ru%3A60889%2Fannounce&tr=udp%3A%2F%2Ftracker.cyberia.is%3A6969%2Fannounce&tr=udp%3A%2F%2Fexodus.desync.com%3A6969%2Fannounce&tr=http%3A%2F%2Fexplodie.org%3A6969%2Fannounce&tr=udp%3A%2F%2Ftracker3.itzmx.com%3A6961%2Fannounce&tr=http%3A%2F%2Ftracker1.itzmx.com%3A8080%2Fannounce&tr=udp%3A%2F%2Fp4p.arenabg.ch%3A1337%2Fannounce&tr=udp%3A%2F%2Fopen.stealth.si%3A80%2Fannounce&tr=udp%3A%2F%2Fwww.torrent.eu.org%3A451%2Fannounce&tr=udp%3A%2F%2Ftracker.torrent.eu.org%3A451%2Fannounce&tr=udp%3A%2F%2Fretracker.lanta-net.ru%3A2710%2Fannounce&tr=udp%3A%2F%2Ftracker.ds.is%3A6969%2Fannounce&tr=udp%3A%2F%2Ftracker4.itzmx.com%3A2710%2Fannounce&tr=udp%3A%2F%2Ftracker.moeking.me%3A6969%2Fannounce&tr=udp%3A%2F%2Ftracker.tiny-vps.com%3A6969%2Fannounce&tr=udp%3A%2F%2Ftracker.zerobytes.xyz%3A1337%2Fannounce
If the current push of AI companies to get their way (to allow copyright laundering) succeeds, this would almost count as open source by the real definition.
If not ... lots of people/companies are committing copyright crimes, some are committing civil infractions, and some may be able to claim fair use.
[1] https://github.com/Hellisotherpeople/DebateSum [2] https://github.com/EleutherAI/the-pile/issues/56
Stella was waiting for you to submit your dataset. Did you? She closed the ticket many months later.
Also this was before most datasets were hosted conveniently on huggingface.
It's all tears in the rain now.
The era of intellectual "property" is over. Let's be at peace with that and just move on into the next age.