Now is the time to give LLMs access to the ACM digital library
cacm.acm.org
cacm.acm.org
I would be surprised if a majority of ACM members were to say yes should we ask them (but ACM is not known for such democracy). Along with book authors, we are one of the many people that provide the knowledge and expertise on which large tech firms train their models, and get nothing in return. Actually, life is getting worse for us: extra workload in universities with students' AI use, a completely broken peer review system, etc. Hence the irony of ACM thinking about licensing, and only licensing, at a time where this is the least of our priorities.
You don't own the knowledge you put out there unless you have a limited time valid patent. The rest is absurdity. If you want to keep your findings to yourself, keep them secret.
The entire point of copyright law is so that people can make their writing public and still be able to control the right to make copies (for example, into your dataset for training an LLM).
The output of it is also a derivative work, and derivative works also infringe copyright. Its only not a problem if you ignore copyright entirely
Humans are the only entities that get to enjoy special idea-learning-exemptions, not AI
Derivative work or transformative? It's not the same.
Not only are you not winning this one but I'm gonna laugh at you the entire time.
Kadrey v. Meta Platforms, Inc., No. 23-cv-03417 (N.D. Cal. June 25, 2025)
Fair use is a defence against copyright infringement. Ie you actively say that you *have* committed copyright infringement, but you're allowed to do it under fair use doctrine to train the model. That says nothing about the purposes the model is used for
There's also these parts:
> its use of pirated books to create such library does not constitute fair use.
Which indicates that there are tight bounds depending on the ethics of how the content was obtained
Similarly with the second one
>Meta moved to dismiss plaintiffs’ cause of action for direct copyright infringement only to the extent that it was premised on a theory that the software comprising LLaMA is itself an infringing derivative work.
We're talking specifically about the output of the models being infringing, not whether or not the models themselves are infringing. If you read onwards
>Plaintiffs’ claim for vicarious copyright infringement failed because the complaint did not allege that any output generated by LLaMA contained protectable expression that recast, transformed or adapted the books. Without “an infringing output, there can be no vicarious infringement.”
Which strongly indicates the precise opposite of what you're saying, if you actually like, read the rulings
I just pointed out none took place.
Anyway I'm not replying in this thread anymore.
> and yes that's 100% copyright infringement
Says what court of law?
I'm kinda getting tired of this stuff. I'm someone who has been, and still to some extent is, uncomfortable with the possibility of copyright/license laundering in LLMs, but they way you are making your argument is incredibly off-putting and not sympathetic. You're throwing out wild assertions about the law that are not supported by... anything, really.
The main question in the "AI image generator generates a Pikachu image" is whether the AI company serving that image generator to you is violating the copyright or not. Because they make money when doing so (API / subscription cost), and so it's like selling images of Pikachu. The user is likely in the clear as long as they don't go on sell that Pikachu further. But the AI company sold the Pikachu image to the user.
>Says what court of law?
If you turn a png into a jpeg, and distribute it, that's copyright infringement. There isn't a court in the land that wouldn't find you guilty of that
Only every movie piracy lawsuit ever. Nobody shares the original files after all, so every torrent is re-encoded in the way described.
AI works also clearly aren't transformative in many cases. If you ask it a question about a paper, it'll quote bits of the paper at you. That serves as an exact substitute of the original work. If you ask it for song lyrics, or information about the news, its content is a direct substitute for the original source it was trained on. This clearly does not fall under a transformative use case
You could argue that some uses of it are transformative, but even then - its easy to find some piece of training data in the source code that the output work supersedes. By its very nature it does not have the capacity to genuinely invent under the law (as it is not human), and a prompt isn't a significant enough part of the processing to count here
Everyone treats the human user of the AI as furniture, but they steer the whole process into unique directions.
You can use all the ideas you want, but AI cannot because its not a person, and does not enjoy the same protection under the law. The copyright holders by and large did not agree to you using their content like this
If we enable this, people won't create anything because all their work will immediately be stolen by the AI models. Copyright partially exists to promote the creation of new content, because theft disincentivises novel creation
1. AI models frequently output large chunks of code which are plagiarised. In one specific case it was code for walking the stack, that was a clear mix of two original sources that I was able to find with changed variable names, but the structure was identical and switched from the first to the second halfway through
2. AI models plagiarising stack overflow answers word for word, quite recently about the rotation rate of smoothbore cannons in the age of sail
3. AI misspelling answers because the physics papers its trained on made the same typos, which is how I discovered that it had plagiarised the answer
4. Misconceptions/wrong answers that can be traced back to specific papers due to the oddly specific nature of the language used
There's been a lot of research about getting AI models to output their training data, and it turns out they store huge amounts of it. You can use this to get people's personal information if you really want to, and that's very low occurance information
1. The plagiarism aspect, and that most of the training data was used without permission
2. I haven't found it terribly useful in my personal work, as the data it was trained on was heavily polluted by incorrect information (at least in the field I'm using it)
Also, the restriction isn’t on a technology that could possibly reproduce something. It is on the act of using the technology to reproduce something.
A short session is just retrieval, a long session is always unique. The more the user writes the more it diverges from any content in the dataset.
As someone with lots of open source code out there that has likely been used as LLM training data, I'm very sympathetic to this point of view, but that doesn't seem to be the legal reality. Much of this has not been fully tested in court, but it seems likely that LLM training is not copyright infringement, as long as the training material itself was acquired legally.
There's also been court cases where material has been found to be infringingly used, eg song lyrics, so the case where copyright ceases to exist doesn't seem to be coming through yet, thankfully. It'd be the most staggering upheaval of copyright of all time if this doesn't turn out to be true
I am not personally affected because I don’t mind LLMs using my code and writing to learn. I have open source under MIT and similar licenses. I didn’t foresee LLMs learning from it, but it does feel like it’s in the spirit of what I intended.
The issue is not who owns knowledge, it's how it benefits humanity.
If someone reads a philosophical essay, has a heureka moment from it, applies the principle to their work, and makes bank (commercial profit), they never have to pay a percentage to the author of the essay.
Currently patents are predominantly used to prevent interoperability and impose costs.
Ha. Are you new here?
The billionaire are not going to magically develop a sense of empathy and morality and give away huge chunks of their fortune to fund Universal Basic Income. They would rather invest a few million into building even taller walls around their compounds to prevent a repeat of the French Revolution.
Sorry to ruin your day, but if the people with money could have introduce equal society, you would have noticed their attempts by now.
The take of "if there would be no resource distribution problem, there would be no problems" isn't a genius take, it is literally a ancient children's story. Granted it is factually correct: If there were infinite resources and they all would be in the right places, there would be zero problems based on a lack of resources.
However the problem with this theory is everything else, the first question being:
How do you get from a system where many of the most powerful actors profit from (often: artificial!) scarcity to them giving up those power levers voluntarily?
Or phrased much less abstract: Few bosses, want to stop being above others if they have the choice. And not only do they have the choice, the mechanism some people believe would lead to abundance (AI) is in their control. Turns out the computation power possible inside a finite universe, on a finite planet, with finite resources, an finite labor is.. well.. finite.
Aside from that we already live in a world where the problem is nearly never the amount of resources, but the way they are distributed. Clothes that are produced for the west are shredded after each season as they go out of fashion, while in other parts of the world people live without clothes, with food the story is similar. The problem isn't lack of abundance. The problem is that a wealthy minority profits wildly from the inequality and if you ask me, they would rather use AI to make it stay that way, while extracting even more from the rest of us, than using that for good. Why do I believe that? Well I do not trust them when they say what they want, but instead look at what they do. Just look at the history of capitalism and where within it truly wealthy people started to contribute to society. You may be surprised if you look at the deeper causes...
This theory of abundance is a story followers of an ideology tell themselves so they can feel good about themselves, while their ideology literally starves people and burns down the planet.
Abundance, by it's definition would not only mean everyone, but everywhere. As I argued in my other reply, I believe that this everywhere is the much bigger challenge than the everyone.
It's a racket.
Guess the single strongest predictor of paper in field per year?
but you want sweet grant money, you play the game by the rules.
researcher that don't want to play the grant game can go and find a VC to subsidy product creation
A non profit could well pay, and there are plenty of reasons frontier models should be managed by non profit. After all, why allow a for profit to benefit from free contributions?
If you're offering me $300/hr and the other guy is offering me $400/hr, a whole lot of things start to matter more than the differential.
That gets at the "fungibility" notion you refer to. Money has the same nominal value everywhere, but different real value based on how much of it you already have. Which is to say that the marginal value of money drops pretty precipitously several times at certain thresholds that relate to the particulars of the economy (when you can afford to eat, when you can afford a house, when you can afford to not work anymore, and so on).
I’m not sure which Catholic saint said that, but one requires a minimum of material conditions in order to practice spiritual virtues.
> when you can afford to eat, when you can afford a house, when you can afford to not work anymore, and so on
When you can buy a government, a couple billion seems like a small loss.
And when the open weights models distill all the content out of the majors anyway?
Yeah, too bad
I get the copyright aspect of this and I'm not arguing that here. I'm more asking about the moral / ethical idea of choosing who can benefit from your science.
Obviously there are the moral / ethical arguments about AI in general here to weigh against - those have been hashed out significantly elsewhere, and I'm not interested in debating them. What I'm asking about here is the impact on science by sharing it with tooling that distributes it in ways not generally considered when originally written.
A quick check of your post history suggests the frame that you work in strongly is privacy related research (observation - may be wrong). I'm curious how that impacts what you wrote here generally.
(Just to be perfectly clear, I'm not arguing your points here, trying to understand them better)
That said, what matters here is the social contract, what do I bring to society and what do we get from tech companies. For most people around the world, access to the typical leading models is out of reach. Not many on this planet can pay the subscriptions (or even API keys) that offer access to the best models. So I'm not buying the argument that tech companies are broadening access. What we're creating is a increasingly discriminatory society where the few get access to information, and the many don't.
I guess I have some perspectives on a bunch of this. I'm for open sharing of academic work for all (but I'm not an academic, so my perspective is a consumer), so inferring your perspective here I think we agree on that. I maintain many open source (MIT/Apache2 licensed) libraries, and I've also worked in big tech (Amazon, OpenAI). I believe both in the idea of collective commons but also in the ideas that there should be the ability of people to sell software. There's tension in that social contract similarly, and it gets more complex when you look at copyleft.
I guess I'd be disappointed if this was just allowing big labs access and not more broadly allowing access to the ACM library for smaller open source models. Very much in agreement with your last points there.
I’m sorry, but I can’t buy this argument. Making information more available does not make it more discriminatory. Nobody is saying it will only be available in the best models and withheld from other models or services like the ChatGPT free plan. Nobody is saying we’re going to make the original content inaccessible through the previous means after the LLMs are trained on it. Nothing about this shrinks access or makes it more discriminatory.
I understand that you’re upset about the use of the content, but I think you need to admit that your stance is the one trying to restrain use of the content. Training LLMs on it can only bring knowledge to a wider audience, not restrict it.
Whether or not that’s a good or fair idea is a separate discussion, but arguing that this makes access to the knowledge more discriminatory and locked away is 180 degrees backwards.
Is that not why they’re shredding the books when they’re done with them?
From here:
https://x.com/HedgieMarkets/status/2081534588485296565
A quote:
A federal judge ruled the practice is fair use because eliminating the original means only one copy exists at a time.
So no, the reason isn't to keep that information, even from competitors. It's about a legal ruling, which allows them to scan said books without infringing copyright. Judicial decisions and case law have lead to this outcome.
Please stop spreading rumours without any validity.
Gotta love that ridiculous case law system :D
As the other commenter posted, the original report that AI companies were shredding books was based on a second-hand retelling of a rumor, embellished for "AI bad" headlines.
There are actual bookshops talking about this are saying that most of the books are things like "How to master Microsoft Word 96" and that's why they're rare. They're also saying that the destination shipment is going to FBA (fulfilled by Amazon). They think it's an flipping operation trying to find arbitrage opportunities.
Also the reason AI companies have to destroy books is because they've been legally forbidden from using digital copies available. They had to pay a large settlement for it. So it's not some conspiracy to deprive the world of knowledge. It's what the courts told them they had to do.
https://en.wikipedia.org/wiki/Paradox_of_tolerance
Without material values, you will be lost and confused.
Granted, the ACM is hardly a bastion of anything but self interest and greed.
It is good for their business model; but it may lead to monopoly unlike anything acm ever had and very dubious prospect for academia.
The OP didn't say that.
What am I missing? Why should we be concerned about AI leading to a monopoly? And for which company?
Or did I miss the change, and ChatGPT, Claude and Gemini no longer have a free tier anymore? And then the next tier that costs peanuts for anyone in the west, that gives you more access to slightly fresher models?
The first ones are free.
And as long as they keep the weights proprietary (which all the major players are for their primary models), they can decide to charge whatever they want later.
Completely false. The price for LLM inference at a given level of model intelligence has been dropping like a rock.
There are more expensive models available, but you don't have to use them. The same LLM knowledge that was available a couple years ago is now free to download and run on your laptop.
You’re talking about companies using more tokens.
Completely different concepts. As I said, the price for a given quality of LLM output continues to decline. Has nothing to do with companies using more tokens.
The peer review system was broken before AI, so I’m not sure what your point is there.
The copyright agreement (actually copyright assignment/transfer) with ACM has had a carve-out for this case in it for a while, but it has to be a personal website.
https://dl.acm.org/openaccess ---> So how to access the content? Do I have to register or what? It the "open access" only for academic org's people or for everyone in the world?
https://dl.acm.org/ ---> Okay, nice simple search field without loggin in, but when you try to search something, you get thousands of results, even if you search specific author and the exact name of paper, you will get hundreds of results and the thing you want is buried somewhere on page 247. Filters of authors, years etc. are for Premium subscription. But just googling the thing finds the link to ACM... And want to get the actual PDF? Hope it says "free access"...
Google scholar has become the way to search papers (which is somewhat worrisome). What people want from there is a download for all the stuff that does not list an author copy. Open access IMHO is just a reaction to the fact that mostly you would not need a subscription anyhow. Now the authors are paying upfront or universities are paying flat for all their researchers. The problem is now the incentives are not increasing the number of subscriptions but increasing the number of papers published.
https://authors.acm.org/open-access/acm-open-for-authors-hom...
The back catalog at ACM isn't subject to this requirement and nor is research that's not funded by the US government, but I completely agree with the spirt of this law: research should be shared knowledge that other intelligences can build upon, whether human or machine. If you want to limit access to what you've done, don't publish it. Get a patent if that's an option, or keep it internal to a company as a trade secret.
But no, it will presumably get much worse as LLMs are inserted into this equation as yet another and new gatekeeper.
(Disclaimer: I have publications with ACM, non open-access. And ACM wasn't even too bad, it's the others that give me pause.)
I think the right choice is pretty clear...
Especially in mathematics, many specialized areas have fewer than 100 people worldwide who are capable of determining whether a result is correct or not. The mathematics community will certainly be willing to use AI to assist with their research, but they may strongly object to a flood of AI-generated mathematics papers produced by others.
What about the authors?
I suppose they could be sued.
There's something hilarious about that, but also, snake eating its own tail.
A llm emulating a person is why many of my used sites banned llms due to scraping bandwith costs
Is this about access or accessibility to claude (for example)
I am sure the entirely of human computing knowledge is not that big.
Does the ACM really think LLMs haven't already consumed 90% of the content through other sources?
Either way, I'll pirate.
now it will be fodder for the slop machine
(i think that LLMs are going to wreck the peer review system for all but hard-experimental papers)
> Beginning January 2026, all ACM publications and related artifacts in the ACM Digital Library will be made open access.
I think I may be shadowbanned, but at what point do we start viewing LLMs as a national security threat?
AIs, at least in their current form, make you more who ever you were. If you want snap, glib answers of dubious accuracy, they'll give them to you, more easily than ever before. If you want to dig back into primary sources and get the original content, they'll do that for you, more easily than ever before.
Can't speak to how the science infrastructure is going to handle them, but if it takes down the peer review system, which I think has been worthless for probably going on two decades and has just given the entire enterprise a false sense of assurance, it'll probably be a net gain in the end. Peer review is a source of more problems then it is solving right now.
It's almost like, "don't hurt yourself unc, we will just search arxiv".