After all, China already blocks the Western businesses that don't implement the spying/surveillance measures the Chinese communist party wants.
The more the Chinese communist government continues pushing a war course against Taiwan makes that even more likely, e.g. in the context of sanctions.
If AI made all the art that might sell, that would give me more time and energy to work on the art I actually care about, but maybe less sales from art.
The idea that artists should "do stuff for the joy of creating" is just plain insulting, too. They already do that. While artists would love to see, say, AI art models that were trained with licensed or public-domain data; training data theft isn't even their biggest concern. Their biggest concern is having the fun sucked out of their job as the artful minutae of drawing or writing is replaced with finding the correct combination of words to make Stable Diffusion draw the character you want with exactly the same details every time. It would be like if you worked at a PC building shop and one day the boss said "Actually we're just going to be an Apple authorized reseller now." The thing that's destroying artists' jobs being trained on their own work is just insult to injury.
To be clear, though, AI doesn't "steal and resell content" in the vast majority of cases, either. Regurgitation is a thing, but the cause is duplicate data in the training set making it advantageous to memorize a few images to improve loss metrics. Most diffusion model architectures are not big enough to memorize the whole training set, or even large pieces of it.
So no, I don't consider that insulting.
If only our law makers were so capable of a) understanding whats going on and b) actually passing laws when it mattered.
[1] - https://arstechnica.com/information-technology/2022/12/china...
Are they going to enforce it for products/services sold overseas? I'm not so sure, not until I see watermarks in Tiktok's content.
Like, are you a troll or you just don't know how laws work?
>I'm not so sure, not until I see watermarks in Tiktok's content.
Tiktok literally exports all videos with watermarks, its the people that remove them to post on reddit lmao how clueless are you
About watermarks, I'm obviously talking about AI generated content watermarks, not general tiktok watermarks. I guess I should've been more concrete with my phrasing as some people can't read considering context.
Also, if you follow through your logic you are arguing for abolishing all laws in the US that constrain US companies if some other countries can ignore these laws and get the upper hand. Keep in mind that these laws (like intellectual property, copyright, patents) is the reason US has innovated so much in the first place
After that, they will decimate global gpu supply and they can use their newfound lead in compute to win the race.
Now, if you have any information about AI and China, enlighten me. The more one knows the better.
PS/edit: I need more information about your last ghost edit. China (and others) would say otherwise about IP laws. If anything I would at least say that innovation can be either cut or assured through IP laws depending on the domain/technology, but it is difficult to conclude that absolutely for all cases.
It's not going to make China's AI better than US. Not nearly enough, they're so behind. So giving them competitive advantage is not something that should matter when you consider whether this law would do good or not
on edit: I just realized that of course British copyright law is also very similar to the American model.
Would you not say that using an AI to do essentially the same thing, just with the benefit that you don’t really need to pay anyone for creating the art meaning you can target more niche preferences is also “cashing in”?
Generative AI is going to allow an amazing diversification and explosion of art and music, which is going to create the next big AI application once current systems for distributing content are strained to overloading - interactive recommenders. Imagine asking for a piece of content, getting recommendations, evaluating a short clip, giving feedback to the recommender and getting better recommendations as a result in a cycle until you get exactly what you want.
At least part of the reason I like technical music is because it is challenging to play for the musician. If an AI 'plays' 'technical' music, it loses its appeal for me.
Similarly, a lot of the reason I like metal music is because of the emotion and energy poured into it by the human creator.
Being able to ask a computer make noises which sounds like some musicians you like is not, in my opinion, the same as creating art.
Humans have already been able to do this of course, but the difference is scale and automation. With this software one could, under current law, setup a 100% legal 'shadow library' that effectively infringes on every single book published. Even a leaked copy of a book could be released and shared (in its legally infringing format) before the "real" copy hit the market. And again, all completely legally. The impacts of copyright infringement are regularly grossly distorted, but I think this is the sort of technology that could genuinely damage artists and creators across many endeavors.
The exact same thing will be coming to software soon enough. 'Take this assembly/IL code, and create a functionally identical but superficially rebranded program while working to ensure you sidestep all relevant patents.' 'Sure, here you go.' The issue of this being done [relatively] instantly is going to really impact things in ways I think many are not considering. Never in a million years thought I'd see myself on the side of the copyright cartel, but this is one of the extremely rare times they're right.
The barrier to shadow libraries is marketing, which is the only thing that really separates huge blockbuster music/art/books from stuff that makes literally zero money. Quality and ideas haven't been the gate for a long time.
If not a single new piece of art were created from this point on, I'd still die with a massive backlog of material to read, see, and listen to.
At this point, I'm infinitely more concerned about the wellbeing of artists than about whether or not new art is going to be created.
> is super encouraging to the artist.
This is the "nobody goes there anymore. It's too crowded!" version of making art.
Creating more, or "limitless" art is indeed the point.
But neither of those is really what's written down, with the (U.S. law) point of copyright being "to promote the Progress of Science and useful Arts".
We might could interpret this as, training AI is progress of science, and thus the point of copyright includes training AI. Although I'm not sure that we really need copyright to do that, and accordingly, I doubt that is the intent of the law.
Really, the issue here is may be that copyright law was not formed with AI systems in mind at all, neither as creators nor as consumers, and trying to apply it to AI systems, or to reason about copyright law as it pertains to what AI systems do, doesn't necessarily work very well.
Maybe we need to go back and amend all of our written laws with a phrase like "for humans", just like so many science headlines need to include the phrase "in mice"!
if only Microsoft (OpenAI) are able to exploit works for material gain at the cost of the literally 100% of the rest of society: why should society allow Microsoft to do this?
(or even to allow Microsoft to exist at all?)
We may find that letting content creators choose to have their work included, or not, in AI training sets would be helpful? Or, included in training sets for a fee, or for some sort of attribution, or ?...
The apparent current status quo of "anything on the public web is fair game for an AI training set" might not be a good permanent solution?
IMO, what LLMs are demonstrating are the fundamental contradictions of Capitalism, where now being a capital owner with a bunch of GPUs is supposed to give you exclusive returns on the sum total of human intellectual and artistic labor. I have a feeling that people aren’t going to take to that too kindly, so we’ll either see robots mowing down the masses of unemployed, a Butlerian jihad, or various states assuming control over their productive capacities and redistributing the return in the form of greater safety nets.
Big platforms really are not helpful nor respectful to musicians, especially musicians that are working hard to be discovered. It's a total shame that shotify charges musicians to be promoted on their platform while giving a ton of royalties away to so many fraudulent actors every year.
I don't think that would work, if a song is detectable enough to give the artist royalties for a song, it's detectable enough to see that it's copyrighted music. IMHO I think music piracy is essentially dead, because the music industry has finally learned and made a compelling product. With services like Spotify and Apple Music, pirating music is just not worth it anymore. Why bother when for $10/month you could listen to all the music to your hearts content.
Universal asking steaming companies to not use their label's music on training is just stupid because all they are doing is shooting the artists in the foot. Most AI in music is recommendation engines, do you not want your artists to be discovered? The streaming services are not stealing your music you idiot [UMG], they are trying to make you more money by directing people to music they like.
Ai generators put music and other content in a blender and then scrambles the source samples until they are unrecognizable and then mashes the tiny cut pieces of source samples back together into a collage. Kind of like putting a strawberry into a smoothie... It's no longer a strawberry, but the smoothie now has the extracted taste of strawberry AND the other original materials used to source the end product. Content ID only recognizes strawberries and whole fruits, not smoothies.
Map-makers use arificial "trap streets" ( https://en.m.wikipedia.org/wiki/Trap_street ) to show that their maps were copied. Dictionaries add fictious entries ( https://en.m.wikipedia.org/wiki/Fictitious_entry ).
Similar idea could be used with music?
the process the work was generated through is of primary importance for copyright
the existence of the concept of "derivative work" should make this obvious
see the reverse engineering of the IBM BIOS for another example
we're not talking about fair use here buddy
your honor, this LLM produces content that looks like various Disney intellectual products, it was obviously trained on that data.
your honor, no we trained it on a lot artwork that was copying the Disney style!
Judge: do you have this corpus of data you trained it on for the court to inspect.
uh, no.
summary judgement for plaintiff.
Serious question.
Since the hypothetical AI company wasn't able to produce anything in it's defense, it lost.
IANAL
Strictly speaking, you do have to prove things (if you have the burden of proof on the specific issue), and the standard of proof is usually “preponderance of the evidence” (though there are a few other standards that apply to particular issues/circumstances in the civil justice system.) “Beyond a reasonable doubt” is a different standard of proof used for conviction in the criminal justice system, but it is strictly not correct to call proof under other standards something other than proof.
I've edited my comment to reflect your correction.
I don't think the suggestion is meant to be that this is a realistic scenario though, just that the court isn't as mechanical and naive as we programmers are prone to imagining (being stewards of systems that are largely mechanical and naive).
IANAL
so after the several days when Disney shows how this LLM was obviously trained on the dataset of available Disney content, and then the defendant responded that they trained on a corpus of pseudo-Disney but then could not produce this corpus it would be reasonable to conclude they were lying.
The parent suggested "No one will be able to prove the source of the data" and that is the kind of thing that programmers for some reason often think is some fantastic gotcha so one can really get away with anything, but it is these kinds of things that the law is generally pretty good in handling.
"Clean Room Defeats Software Infringement Claim in U.S. Federal Court" http://hoviblog.blogspot.de/2008/10/clean-room-defeats-softw...
"Chinese Wall" https://en.wikipedia.org/wiki/Chinese_wall
This has nothing to do with gzip.
[1] https://www.federalregister.gov/documents/2023/03/16/2023-05...
jpeg'ing frames of disney animation doesn't remove copyright (even if you do it twice)
I can write an algorithm that just generates random noise in the dimensions of art, and if I run it long enough it'll output things that are "close enough" to copyright works. There's no argument for that program being copyright infringement, and that holds for models as well.
right, so if it's a completely original work let's clear out the training set and let's see if it can do it without it
no? so the output is a product of the input... a derivative work
and if you run it again it produces exactly the same thing? (sans artificial random injection)
sounds like a lossy compression function to me!
Cuz I don’t.
Training on shite results in a pretty poor AI compared to one that’s trained on quality data. Do I think that means that quality AI will keep up with “hoover AI”? No. Hoover AI still gets you to profitability and that’s mostly what matters in a capitalist arms race.
"A Child left in an empty room is never going to get ahead of one that has experiences with everything in the world"
Anyone that thinks a clean room is going to make a valid AI, in my mind, is insane.
Stable Diffusion and the Getty Images lawsuit will end with a settlement and licensing will be the option to go with.
There's a cool song called 'Neural Harmony.' The creators used a mixture of classical and electronic music as input, but they never disclosed the specific songs they were influenced by. This has made it impossible for the original artists to claim royalties, and yet the public can't get enough of its sick beatz.
There's even an AI-generated album called 'Digital Renaissance.' The creators claim to have used thousands of songs from various genres as inspiration, but they never provided a list of the songs or artists. The album has gained a massive following and has even been featured in several popular playlists. There was briefly an attempt to prosecute them but the case was dropped after public outcry.
quite well then?
streaming services have essentially replaced music and video piracy
and steam (essentially streaming games) vastly reduced video game piracy
Gabe Newell once famously stated, "Piracy is almost always a service problem and not a pricing problem." Steam's success can be attributed to addressing the service issues that initially led people to piracy, providing a user-friendly platform for gamers. So, while DRM might have had a minor role in curbing piracy, it's the innovative business models and improved services that made the most significant impact. Even now, downloading movies, albums, or video games for next to nothing remains possible, and those who prioritize money over time continue to pirate without issue.
who's going to bother for music when spotify/youtube are free?
your time would have to have negative value for it to be worth it
So I imagine, someone will train a model to do high end law work and they will have to train it on the data produced by the people they hire to create it.
This of course assumes that the society will have the same structure as of today. What I actually think it will happen is, we will completely delegate all our work to machines and the concepts of ownership will vanish as everything for exception of land will be in abundance. I imagine in the future we will fight each other over apartments in cool areas or trade it for some kind of social credit which we generate by impressing other humans. No apartment will be worse than the other but the proximity natural wonders, cultural centres and networks will be the paramount. After all, it's all about how we pick our sexual mates and social status.
I'm sorry you don't like it but There's no way the society functions the same once resources and servants are in abundance. Your bank account is relevant only when there's a scarcity and the current scarcity can be gone once machines are autonomous in enough areas to convert the material all around us into things we need using the practically limitless energy from the sun.