The I Hate AI License
ihateailicense.eu
ihateailicense.eu
Those flaws probably exist because the text was written by AI, per the footer - "This content is fully generated by AI because fuck AI."
I wouldn’t frame this as stupid, or even trolling. An art project maybe, and one that is doing its job seemingly pretty well.
Laws are written based on the results of these kinds of discussions. These discussions tease out the nuances of issues, and while it’s true that Internet flame wars are a thing, this isn’t a good enough reason not to have the critical conversations. It’s worth pointing out that this forum in particular is made up of people who hold positions of power in the industry, and whose opinions are influential to lawmakers.
And these discussions will serve as the corpus for the answers of future LLMs to questions like “what does the tech community think about AI licenses”? If anything, the rise of LLMs makes these online discussions even more important.
There is dialogue, and there is physical conflict. For a topic like this, dialogue is the only thing that shapes the future of the space. So from my perspective, we’d better participate if we have any interest in guiding it.
Do you think that an average anti-AI internet activist would do a better job all on his own?
I'm also curious if the usual way of reading web content includes an "AI" somewhere in the path (ie: browser, upscaling, fraud filtering, video compression, CDN, etc.) between the server and the eyeball.
I think that depends on the application of "designed to simulate human intelligence" (Definition "b").. If that statement does not apply to the previously mentioned "algorithms" in the definition, I think it's safe for most (current) search engines.
I do believe the license would prohibit you from passing this through an intermediate AI system, as there's no exceptions or allowance.
I imagine there would be a fight in various laws/constitutions saying that the license is illegal/not allowed/etc in most places as it would not support accessibility (e.g. AI based text-to-speech for those that need it)...
And then there's no severance clause/statement, which is technically in the CC licenses (legal version), to protect the rest of the license should one part be found invalid.. so ultimately the whole license is probably invalid for most countries since the limitation will almost certainly (eventually) violate someones rights to read the public information (which will probably have AI incorporation at some point, e.g. text-to-speech with automatic summarization, searching by context, spelling corrections, perhaps translations, etc.)
But then, what does it apply to? Are LLMs "designed to simulate human intelligence"? What about image generators? Like, their outputs can be similar to what "human intelligence" might produce, but if we're talking intent at the design level, I'm doubtful that any of their developers sat down with the thought of "let's replicate a human".
> I imagine there would be a fight in various laws/constitutions saying that the license is illegal/not allowed/etc in most places as it would not support accessibility (e.g. AI based text-to-speech for those that need it)...
Either this or "right to public info" you mention in the next paragraph are smaller issues in comparison to asking whether they can even make these clauses. If AI use isn't found to be automatically in violation of copyright just through mere inclusion in a dataset, then this license will be meaningless because you can't sign away law in a license.
>And then there's no severance clause/statement
Probably this a point on why I'm a programmer and not a lawyer, but I never understood how these could work. If not having them means that one part of a license could invalidate the whole license, then how does having this change things? Something invalidating the rest of the license would also invalidate this part of the license, meaning it would be just as enforceable as other parts of the contract. If this one clause doesn't get invalidated, then why would the other clauses also be invalidated?
(I am not your lawyer, this is not legal advice.)
https://www.centennialofflight.net/essay/Aerospace/WWi/Aero5...
> The United States did not produce any aircraft of its own design for use at the front during World War I.
If it happens again for AI who knows what will happen.
Besides, rather than vacuuming up exisiting data, future resources might be better spent on creating bespoke datasets with user consent and content licensing baked in from the start.
Not mutually exclusive.
We have conducted large statistical surveys in the past, so I see little reason why these and other established data collection practices could not be updated and utilised for refining contemporary models.
It's probably a large reason why this space is so difficult to license/regulate: data to one may simply be 'noise' to another, or not even register as such. Definitions unfortunately, do not always map.
And I think there are dozens of unfinished AI copyright infringement lawsuits... with likely more to come, I speculate.
https://chatgptiseatingtheworld.com/2023/12/27/master-list-o...
https://heathermeeker.com/ai-copyright-et-al-litigation-scor...
> Ultimately, Judge Bibas denied summary judgment and ruled the case should proceed to trial.
> Judge Bibas cautioned against “overreading” the Supreme Court’s recent fair use decision in Andy Warhol Foundation, in which the Court limited the weight to give a defendant’s transformative purpose if it is commercial and substantially similar in purpose as the plaintiff’s work.
> "Ross’s uses were undoubtedly commercial. And one of its goals was to compete with Westlaw. Thomson Reuters contends that this commercial use weighs heavily against finding fair use. [...]"
All that the judge is saying is that previous rulings are not comparable to AI related ones. It is also difficult to see how ChatGPT is competing with NYT, as ChatGPT does not generate (actual, real) news, because it cannot interact with the world as journalists can. Maybe image generators could but a similar judgment stated that it was difficult to see how AI generated images were similar to an existing artist's.
If I'm interested in Dick Cheney shooting someone by accident, and previously I would have googled that and been sent to NYT reporting on the incident (and get shown some ads or sold a subscription) but instead I ask ChatGPT - that seems like a competing product to me.
And fair use is a multi-pronged test - ChatGPT's use of the data is clearly of a commercial nature, and the portion of the copyrighted work used for training is presumably 100%. So while Wikipedia can quote an NYT article (being both noncommercial and a short excerpt) the court could well consider LLM training more in the nature of music sampling, where at times it's been held that even minuscule two second excerpts, distorted and mixed in with lots of other music, still need to be licensed.
I may think TikTok is vastly inferior to the National Geographic TV, but they're still competing for eyeballs and ad sales.
Otherwise you could argue that if I copy and paste the NYT articles onto my own website, because I'm not doing any journalism, I'm not competing with the NYT - which is clearly ludicrous.
Why? NYT paid for and put work into their writing, and they do make some money on holding the archive of their work. Not everything they publish is breaking current events either. Chatgpt does not have to come up with new news in order to compete with NYT archive, their editors, their writing style, or their content, among other possibilities. AI’s very purpose is to absorb the content and style of its training data and be able to reproduce it more cheaply than the original. It doesn’t seem difficult at all to imagine many ways this does direct and indirect harm to NYT, or anyone who had to put effort or money into creating something valuable.
For commercial use, too. For EU though, commercial use needs to respect opt-out. Other countries can ignore opt-outs.
Those licenses aren’t very open. Like the Creative Commons non commercial licenses.
First off, telling someone how they can use something is what a license is. Even the GPL is defined in those terms, and moreover, the GPL tells you how you can use the code it protects! ("If you use this code, you must release the source" is functionally as much of a constraint as "if you use this work, it can't be for profit", the difference is ideological, not categorical)
The only function of bringing that up is making weird black-and-white absolutist arguments where "any license which does not grant every possible use is not open" arguments. That just doesn't match with reality - all variants of CC are much more open than standard copyright.
Also, an extremely important point that is often overlooked in these arguments: your statement is actually false, since it does not tell someone how they can use a work, only how they can use it under that license. In other words, if you want to profit off of someone else's work, you're still free to contact the author and receive license to do so.
As I pointed out, the fundamental premise of the GPL is not "there can be no restrictions in license to use a work", it's "the only acceptable restriction is restricting others from restricting their own work". Which, itself, is a laudable goal, but it's an argument that fs folks instinctively shy away from - because they know it invites the immediate counter, "if you admit there are acceptable exceptions, why is that the only one?"
And that's the rub, because the next question is, "So why is it wrong for someone to want to see the fruit of their own labor when someone else is profiting off of it"? Especially when they themselves are offering their own work to the commons, free of charge. If someone is contributing to the freedom of information, they should be able to rest easy knowing it is not being gated by someone else. Of freedom of Use, Redistribution, Study, or Modification, it can only be said to weakly violate "Use", and even then, only when you refuse to compensate the original author. (Art is also unlike software; you can't "use" a painting to manage your payroll in a copyright-prohibited way)
(Also interesting that you bring up the "four freedoms," because one of FDR's four freedoms - those that inspired Stallman's list - is "freedom from want". And living under capitalism, under copyright, one of the chief ways want is addressed is by profiting off selling your own intellectual labor. If you know that you'll get any profits that are generated directly from your work, it makes it much easier to surrender it to the common good.)
You don't make the honest argument, because you know that changing the argument from "no exceptions in the terms of Use" to "only this exception" becomes an argument of "so where do we draw the line", which is much more complicated than a simple black and white.
I believe your counter-argument is "okay, but NC licenses don't really violate Freedom 0, because you are free to run it as long as you pay the author"? Or like, "Freedom 0 is dumb anyway", which then breaks the premises and we can't really discuss much after that.
Therefore, I will only reply to the former. My counter would be "maybe, but in that case you need compulsory licensing, as it currently is, the original author can refuse to let you exercise your Freedom 0 no matter how much money you offer, and therefore it is violated".
Well, one of my points is that's a different argument entirely from the one in either the parent or the GP, which is part of what bothers me. This is the first fully articulated, coherent argument in this thread.
That said, my actual counterargument (which I think should be clear from all of my previous responses) is "well, if that's your argument, then CC-NC is clearly good!".
As I've said from the beginning CC-NC is a much more lenient licence than almost any other. It sets a work to be completely free by three measures, and mostly free by the fourth. By any reasonable definition, it brings a work "closer to the four essential freedoms", your stated goal. It only stops a little short of the goal.
The only world in which this is not true is the you assume that any CC-* license would otherwise be GPL-like. But that's a totally specious assumption - for one thing the default in legal society is full copyright. And most artists wouldn't feel comfortable fully releasing their work for Disney to profit off of while they try and survive off of cup ramen. A world in which CC-NC does not exist is a world that, on the net, I'd assert would be much less free.
This isn't a black-and-white world where you can claim in good faith that "any not-fully free work is just as bad as being fully copyrighted" - and you can't condemn CC-NC unless you're willing to claim that those works are better off not free at all, which is a betrayal of the principles you claim to hold.
> Why does CC support “non-free” licenses at all? Isn’t that against its mission?
> CC hopes to promote a more open culture than “all rights reserved”. Some creators want to allow very broad use of their works. Some wish to allow some uses but restrict others. We believe that all levels of openness are worth encouraging and supporting as an alternative to “all rights reserved.” While we hope that some creators will use the non-free licenses as a stepping stone to greater openness in the future, CC encourages sharing under any of its licenses as a way to create a more open culture.
Source: https://creativecommons.org/public-domain/freeworks/
The only thing that I ask is that certain terms (e.g. Free Cultural Works, Open Source, Libre Software) retain their meaning so there's a clear goal to work towards. One thing that annoys me, for example, is companies releasing software or weights under a non-libre license but trying to pass it off as libre anyway. That's about it x3
(Libre = Free as in freedom in this case, to ease confusion)
Though to be fair GP does say "aside from not sharing it". I interpreted this as "aside from not giving it a CC license". I realise it could be interpreted as "aside from not publishing it" instead. So there might be confusion there.
Broadly speaking, AI encompasses much more than machine learning or deep learning. Wikipedia gives a good view of the goals and tools involved in AI: https://en.wikipedia.org/wiki/Artificial_intelligence
For example, classic search algorithms like BFS and DFS are in fact considered "AI" in AI textbooks, academia and certain industries like games.
If you expand the definition of AI a bit more to include if else statements, then every modern program is probably an AI as well. And rightly so if we view them from the lens of 1990s.
So maybe you want to say machine learning to be more specific.
Maybe this is a language barrier, but what is meant by this sentence in particular? Why "generate with AI" (i.e. an LLM, most certainly not any form of "AI") if the author dislikes that so much? Wouldn't such a deeply held stance motivate someone to more heavily rely on consulting with experts in the fields of licensing and the current state of models rather than outsourcing such a project in its entirety to an LLM? In general, I can understand using an LLM for an initial draft when working on a project, then going over the output with some actual input from various expert sources, though not if one has major objections to the use and existence of LLMs in general. That just seems self-defeating.
Or am I misunderstanding something?
Demonstrating the problems of AI (what it has produced is dodgy in various ways).
A sense of bitter futility, hating AI but realising they can’t actually stop it like this and so it’s not worth the effort of proper drafting.
Perhaps some mixture of these, but mostly the first.
So you believe the LLM was prompted to produce faulty or imperfect licenses on purpose with the intent of highlighting the deficiencies of the current tech, and they did not intend anyone to use these licenses?
That'd actually be an interesting way of protesting, even if it ignores both how prompting can affect LLM output quality and that these models will only improve. Add to that, if one such as myself (and seemingly most people on HN) misses this part of ironic protesting and just takes them seriously (i.e. their intent is that people use these licenses), that harms their cause by making them seem unintentionally unprofessional, rather than intentionally unprofessional.
I don’t read it that way.
You could prompt it perfectly “honestly” for the following:
> create a license for creative works, similar to the Creative Commons licenses, but which disallows using the content in any AI context
OpenAI GPT4 tells me about some considerations, and proceeds to give me a template that I can build on, which is very very near in fact to the kind of content you see in OP link.
No need to tell the AI to create a faulty or imperfect license. The LLM is perfectly able to do that on its own. And then instead of refining what it gave you, set up a basic website hosting the content.
Hell, you can even prompt the LLM to make the HTML version of the document for you, if you want.
> [...] Definitions: Clearly define what constitutes "AI context" to avoid ambiguity. This could include the use of content in machine learning, algorithm training, data scraping, AI-generated art, automated content creation, predictive modeling, etc. [...]
LLMs, by their nature, are not deterministic, so getting differing outputs is no surprise. Considering, however, that, using your prompt, I got a rather differentiated output that points to the very issues that are inherent to these licenses, namely whether things such as data scraping would fall under "technologies designed to simulate human intelligence" (clearly define human intelligence and the simulation thereof, I'd argue LLMs may not fall under that at this point in time either), showcases that the creator most likely did not receive such a faulty license from the get-go, and if they did, the output was likely littered with severe warnings to employ legal advice and lean on experts rather than just post that first draft.
To me, while a LLM may provide such a license draft with such casual prompting in limited cases (perhaps someone could prompt this 100 times and see how often they get a license over general advice using your prompt), the license on that website still seems more like the result of actively prompting for erroneous results to make LLMs appear incapable of such a task (or again and much worse for their cause, this is in fact a serious attempt at a license by someone inexperienced yet unwilling to contact any expert who'd explain things like search engine crawlers and the wider implications of such an undertaking).
Regardless, even if these licenses were the first output the LLM provided and even if they are meant to be taken as a joke/way to start a conversation rather than something people should use for their projects in earnest, what remains is that LLMs are not static and will improve, so any point made concerning the current day's capabilities only undercuts the serious and important discussion on licensing and the future of copyright. LLMs will get better, pointing out temporary imperfections only distracts from actually important topics.
So, whatever their intent and whether this license is serious or not, I feel this does not help their alleged cause.
[0] https://chat.openai.com/share/bf36f182-ba20-4f66-a702-2b636b...
I didn’t want to spam the thread with copy-pasting the full response that GPT-4 gave me for my prompt. But for completeness, here is a link to the conversation:
https://chat.openai.com/share/7295ca81-cff4-4f01-a932-35a3d6...
Archived copy of link:
There you can see the prompt I used (which I also included in my comment above), and the response it gave me.
This was the first and only time I asked it. But as you say, LLMs are inherently using a bit of randomness. So other people using the same prompt will get different variations in the responses.
The act of using "an AI" to generate a license restricting the use of AI can have many motivations, but an obvious one is the simple "use their own weapons against them" principle.
Idk if hn really has a “brand” per se, since it’s just a yconbinator subdomain.
I am not 100% convinced of that, considering we have an almost even split in commentators who believe this to be a serious effort in creating a license and those genuienly taking this to be a more satirical form of joke protest, pointing out the deficiencies of LLMs. Do you think the intent was to create a serious license or poke at LLM deficiencies? I started with the former but have now read the latter so much I am starting to drift...
> How would you phrase this better, in order to prohibit use in n “AI” without effectively banning normal uses?
If we go with this being a serious effort and I had faith in this being a problem that could be addressed via a license, I'd start with researching what I am actually trying to restrict. Copyright law, other licenses with goals of enforcing specific behavior such as GPL and their history, how LLMs are trained and whether this may count as transformative or even fair-use and what the implications of the latter would be if that were to be found to be the case by courts, why content has historically been "crawlable" by search engines and other third-parties, etc. Look at the implications of what I am actually trying to accomplish here and how to go about it.
Then, as a cornerstone of my efforts, explain in detail what the cut-off is for "technologies designed to simulate human intelligence" (what is intelligence, what simulates intelligence, etc.) and, just to name one of the numerous examples, why search engine crawlers would not be covered by that. Or, alternatively, state outright that those would and should be restricted from using any content covered under this license too, and lay out to any potential user what that may mean for them, again keeping in mind the historic background of why website owners have allowed crawlers in the first place.
On top of that, for such an effort as creating a license designed to narrowly restrict very specific use cases, I'd definitely contact and consult with at least a few professionals in the legal, copyright and research spaces. If you approach such a project with a well-thought-out concept and concrete questions to flesh it out, there tends to be a lot of amazing, concrete and highly professional support you can get contacting academia, in my experience. They'll also highlight any aspects that require additional focus and help you address questions before they arise.
I do, however, believe that if the creator of this site had approached some experts with the concept they ended up publishing, they'd quickly be told that restricting the training of "AI" without also disallowing crawling by search engines, etc. in the manner they described is akin to the many decades-long and persistent attempts at manifesting encryption that includes a backdoor, only accessible by "good guys," without harming privacy or safety for the end user.
Very much considered somewhere between incredibly challenging and landing a person on the sun. Due to this, I feel that anyone whose solution is such a poorly thought-out one as these two licenses is either not serious in their efforts or so poorly informed on the topic and its complexities that their efforts won't yield anything positive, not even sparking valuable debate, as such poor ideas always poison discussions, detracting from real solutions.
To stick with the encryption analogy, the continued attempt by some lawmakers to engineer a magical safe black box encryption with LEO-backdoor takes valuable time and resources, as well as public attention, away from actually addressing real ways of fighting those very crimes that are often mentioned when trying to justify weakening encryption for the wider public.
> (Putting aside issues like fair use and enforceability)
I also very much believe any serious discussion can never put these two issues aside. Again, to stay with my encryption analogy, that's like putting "privacy and due process" aside. Any discussion that, from the outset, necessitates ignoring such crucial aspects cannot yield fruitful results.
1. I think whether AI training is fair use is for the courts to decide. In some sense a licence such as this could land a good case to force the courts to do so, as it could prohibit any argument that the training was in accordance to license and as such would require the trainers to rely on fair use exclusively.
2. Enforceability is always a problem with copyleft licenses. You can't expect that putting a restrictive license will completely stop infrigenment, but you create legal peril for anyone who decides to do so.
This blogger has gone on for years about how unfair the Google economy is
It’s like their horses got out 10 years ago and long got caught and taken to the glue factory and they didn’t notice until two magical letters were introduced and now it is fashionable to complain.
It's a social statement, not a true legal document.
For example, if applying some AI-enabled tool to process the data doesn't require permission of the copyright owner (using the copyrighted work without creating copies or derivative works generally doesn't, except in the scenarios explicitly listed in law), then it doesn't matter if the license doesn't permit that, because the user doesn't need to agree to the license.
But it’s not AI that’s making me shift my views. It’s the mess of unresolved problems brought about by the last major generation of tech that gives me pause. Tech that is causing real harms. We still haven’t figured out how to survive the algorithmic social media hellscape we created, and now we’re hurtling towards opening more boxes that would make Pandora proud.
The recent AI explosion has crystallized this view, but did not cause it. My personal view is that tech is nearing some of the very dangerous end points that we’ve only been able to speculate about throughout most of human history.
The atomic bomb completely changed warfare. There are developments past which all old mental models no longer apply. Reducing this to “sides” is a distraction.
I still think technology is part of the solution to the world’s future problems, including the ones brought about by “AI”, but I think it’s foolish to downplay or oversimplify the concerns.
Take the atomic bomb for example: it pretty much ended large scale industrial warfare on the level of the world wars. There are still wars but not ones in which entire continents are destroyed and entire generations across dozens of countries are lost. All things told the nuclear bomb saved the world a lot of suffering.
For the most part I tend to think of social media such as facebook etc as something that will soon be used mostly by out touch old people and hobbyists participating in hobby groups. All things told social media has probably helped the world more than harmed despite enabling Google and others to get filthy rich off private data, in the long term I suspect it will be a flash in the pan caused by the newness of seemless online communication.
Yet. Nuclear weapons are an existential threat and may yet destroy human civilization. We've come very close many times already to starting a nuclear war and we're all basically lucky to be alive. Pour one out for Stanislav Petrov and take a look at https://en.m.wikipedia.org/wiki/Nuclear_close_calls
One of the only reasons MAD has held up is because we've successfully restricted the reach of nuclear weapons, i.e. if every country that wanted them could get them, the average person would have a lot more reasons to be paranoid.
It wasn't always this way. The early Internet was a very different place. I credit some of those early communities with helping me escape a complicated childhood. But that Internet no longer exists, and it hasn't for quite some time. So while I'm willing to agree that social media has provided quite a bit of benefit, I don't believe it continues to do so, nor will it until people categorically reject the current generation of social sites and their business models. This is to say nothing of the privacy and data harvesting issues inherent to these sites in 2024, which by themselves are a threat to democracy when you factor in the overreach we already know is occurring by our governments.
Regarding AI, I think there's a very real concern about social upheaval as entire categories of work are eliminated if not introduced with care. I think it's deeply concerning that fewer and fewer people will prioritize the cultivation of first hand expertise and instead opt to just let various AIs handle things. I think it will lead to the further erosion of independent thought, and to a future where fewer and fewer people know how to reason about things.
The misinformation landscape will be changed forever. People will/are already starting to worship AIs. AIs will be constructed that know how to radicalize people more effectively than any YouTube algorithm ever could. AIs will be implemented as experts (this is already happening now), and we will increasingly defer to these "experts" for critical decisions despite their flaws/biases.
> in the long term I suspect it will be a flash in the pan caused by the newness of seemless online communication
I hope you're right, but this also seems like a very optimistic take. Seamless online communication is here to stay, but we still know very little about the long term impact of this kind of communication ability on humans. The group dynamics in play when our brains evolved look nothing like the group dynamics enabled by this kind of tech.
I should say I don't necessarily think we'll see the worst of every one of these concerns come to fruition, i.e. I'm not just an AI doomer, but I also expect we'll encounter issues we haven't thought about yet, and that the societal impact of even just a subset of these issues could be immense.
Luddites weren't fighting against technology as such, they were fighting against the negative effects it had: factory owners treating workers in ways that would net you many years in prison today.
And it's the same with AI; people generally don't hate the concept as such, they oppose or are concerned with the effects it has or might have. You can disagree with that, that's fine and by all means let's have that discussion! But you can't just brush people aside as "Luddites", as if that's somehow an argument, because it's not.
There's a difference between staying out from under boot and laissez-faire letting a new tech harm your people. The analogy that comes to mind is weapons research: if weapons were regularly tested on civilians to accelerate the R&D, and to enrich huge corporations, yeah, don't call us luddites for being unhappy about it
Eventually, it's going to be the end of humans being able to work in the arts. That's horrific. The most valuable part of the arts isn't consumption, it's the joy of production, subsidized by sales and consumption--and that's going to be gone
Maybe if they trained OpenAI on Microsoft's internal code, or if """Open"""AI released their own source code for free, or anything of the sort maybe less people would be so skeptical, but to anyone without a vested interest in these AI tools it's obvious that they're playing the game of steal everyone else's data without their own data falling into everyone else's hands. After all, if they're so confident these models are safe and won't reproduce Your (the royal You) intellectual property verbatim, why aren't they feeding the Windows internal code to it? Why is it they want everyone else's data, when they have so much of their own?
Many people publish their software under the GPL because they want to enforce that their work remains open and free for everyone. ChatGPT slurping up the source and emitting something based on it, that is then used in a non-free project makes it impossible to guarantee that your work will remain free.
I personally would be ok with my source code being trained on and used as long as the output from the AI is used within the rules of the license I choose.
It would also be a little different if we could train our own "AI", or if ChatGPT and friends were trained on something like Windows source code and the rest of the non-free code on Github. At least then it would be easier to believe that it's actually ok to slurp up code without respecting the license.
The GPL was designed specifically so that you could publish source code without the license being stripped from it by somebody else, that it will remain free no matter who or what uses it or derives from it. Giving the ability to strip the GPL from software as seen fit is something that should be scary and considered very deeply before thinking that it's ok.
They wouldn't, and that's the issue. If I looked at a thousand drawings done in a certain, specific style and used my findings to make my own unique product in that same style, you'd be hard-pressed to find any reasonable people who'd accuse me of plagiarism.
Not to mention that this flavor of "profiting" is something we've been seemingly fine with until specifically generative AI came about. Google profits off of your content directly by using their machine learning algorithms to scrape your website and use its contents in search results, next to their ads. It was found legal, so why should generative AI pay a fee?
Clean room reverse engineering is a thing that exists exactly for this reason but it doesn't seem to apply to "AI".
Also "AI" is not a human, and there is no proof that it "learns like humans". or that it should be even compared to humans in any way, so none of this matters anyways.
It already is. The Windows source code is proprietary data, having unauthorized access to it is already disallowed. You already can't legally use proprietary or pirated data in AI training for the same reasons why you can't use it anywhere else. But this conversation is about the use of publicly posted information that's already out on the internet. It's the equivalent of Microsoft openly publishing the MS-DOS source code and then suing people for viewing it and potentially getting ideas from it.
> Also "AI" is not a human, and there is no proof that it "learns like humans". or that it should be even compared to humans in any way, so none of this matters anyways.
It makes absolutely no sense that a ruling would apply to automation and not humans. If I can legally do a thing myself, can you come up with an argument why I can't build a machine that does that thing for me? The only exception that I can think of is cases where human health and safety is at stake, which is an edge case that doesn't apply to generating text or images.
This is a common fallacy, that having access to the source gives you the legal right to do what you want with it, and that is not correct. The license must be followed regardless of whether it's accessible to the public.
Having access to or viewing Windows source code and pirated content isn't disallowed. The thing that isn't allowed is copying it. That's what copyright laws are about.
The same laws that make it illegal to copy Windows source code are exactly what the GPL are based on, copyright laws.
So why do copyright laws only apply to one side here?
I think there is a lot of confusion and misunderstanding about how copyright and software licenses work here. Windows is able to state that the Windows source code is proprietary because it is their IP, my software is also my IP, so my terms must also be followed in order to legally use my software.
>can't legally use proprietary or pirated data in AI training for the same reasons why you can't use it anywhere else.
This is the same reason why you can't use GPL code in proprietary projects!
The mechanisms that allow something to be proprietary are the same mechanisms that the GPL are based on, this is why it's sometimes called "copyleft".
This isn't what the conversation was about. You started out by drawing comparisons to seeing limited-access information and clean-room reverse engineering, but what you're saying now is drastically different. Do you concede the original point? And yes, just having public access to information doesn't mean that you can do anything with it, but it contradicts the point you made originally.
Besides, the way you describe licensing here makes it sound like the licensee can just ask for anything and everyone will be obligated to follow the rules. This is the crux of the argument, because copyright law has bounds and it only gives authors specific privileges, not a right to dictate anything they like. Seeing past precedent on use cases that are similar to AI, it's not unreasonable to expect restrictions on AI use to not be legally actionable - in the same way how I can't write a license to forbid my hypothetical book from being lent in libraries, or how I can't charge you $100 for reading this comment by adding fine print in the end - a court won't enforce it.
> Having access to or viewing Windows source code and pirated content isn't disallowed. The thing that isn't allowed is copying it. That's what copyright laws are about.
What? No, there are definitely situations in which you don't have to outright copy something to get into trouble. Source code for proprietary software can be legally treated as a trade secret, which gives it a lot of "forbidden knowledge"-type protections that don't require copying.
> my software is also my IP, so my terms must also be followed in order to legally use my software
Again, I'm not arguing against this, there's no double standard in that big corps can have IP, but you can't. Refer to the beginning of my post - the argument isn't whether you can have a license, but whether a licensee can ask people to not use their publicly available work in AI and have it be actionable. This isn't settled law, but you also can't just ask for anything you want in your license and have a court actually uphold it.
> This is the same reason why you can't use GPL code in proprietary projects!
Copying code wholesale from a GPL project is different from having that data be present in a dataset, thereby creating a barely noticeable influence on the production of new content by the algorithm. All this applies very cleanly to direct copying, but with AI, it gets a lot more complicated.
The whole anti-AI self-righteous craze will die down soon enough (as will the "feel the AGI" craze) and Transformer models will be just another tool in the toolbox
Again AI doesnt have any of these hallmarks
This AI wave of which we speak may be slightly different in that it has already proven to solve a certain class of problems, while I'm not sure cryptocurrencies ended up solving any problems in the end, but they both share a similar hallmark of people believing that early promise will lead to complete game changes in the coming years.
That is quite ripe for scams to take hold. It would hardly be the first time where we saw some incremental progress in AI, only to see it stall out and interest wane, but with the grifters trying to latch onto the early hype, convincing others that they have something magical to offer.
It still seems like the whole AI hype is absolutely overblown just as much as the cynical anti-AI hype which I frankly see more of. I really do think it will all die down, we will have another tool in the toolbox and maybe down the road at some point that tool will be instrumental in another hypeworthy breakthrough but whos to say.
It is likely that these tools can make most everyone a programmer, like how improved elevator designs has made everyone an elevator operator. But how game changing is that, really? Tools like Excel have already made great strides in that direction, and with coding now standard fare in public school curriculums, everyone being a programmer is a game already well underway.
Basically - the things you brought up were largely spawned from existing technologies and have very dubious use cases outside of extracting money. AI is neither - it's obviously not a monetary scam as a category, and it has lots of uses by itself.
So, this license bans any use of screen readers and renders the content completely unavailable to people who rely on them.
That said, this includes computer vision (image recognition, object detection, and facial recognition), speech recognition (accessibility).
> You may not use, adapt, modify, or process the Licensed Content in any way with AI technologies. This includes but is not limited to [...], utilizing AI tools...
Of course the other territory we are coming upon is information re use. No existing restrictions are in place. Therefore, rewriting news articles are legal.
(I know someone with hundreds of thousands of words on the web who does want their work to be ‘learned’ by the major LLMs, in the same sense that they want their work indexed by search engines, even if that turns out not to be considered fair use in their jurisdiction, but does not want it copied wholesale.)
So either the license is unnecessary, or irrelevant.