Europe to ChatGPT: disclose your sources
wsj.com
wsj.com
The EU is not anti-AI. In fact, they have stronger protections for AI training than the US does: EU law already has a copyright exception for "text and data mining" (TDM) which covers AI training. The problem is that OpenAI has been incredibly cagey with the way their models get built and trained. This is kind of contrary to the spirit of the TDM exception: it's for scientists to do science with, and OpenAI is being very much not like a scientific organization and more like a commercial enterprise.
"If people are worried one food compagny doesn't respect the laws, they should not use it".
Yeah that's not how law and law enforcement work, AI or not.
There isn't anything left at that point! With that information they could actually have used anything.
If you're going to pretend to be doing science you should at least be held to some of the standards we typically associate with doing science.
I know the article talks about copyright, but not stating any sources for data is a bad precedent to allow.
Woah. Doing bad science isn't illegal, and making it so would be quite chilling. It's common in many fields to be quite imprecise about data used in the work, and entirely uncommon in many for data to be externally reproducible.
Legislation restricting research isn't the right way to improve science and is unlikely to achieve the intended effect for many reasons, including that it's easier and safer to just not touch the impacted area. In some domains this causes whole areas to go unstudied or understudied, e.g. because it runs into IRB and just isn't worth doing... but at least the rules demanding IRB approval are intended to keep people from suffering grave harm and even those are less strong than blanket regulation (they're rules tied to federal funding, not research in the abstract).
What is this "legislation restricting research"? These companies are not doing "science".
I want to know if OpenAI used say GPL or other copyrighted software and then the bastards had the genius idea to put restrictions on the output in their ToS. I want stuff to be fair, if MS/OpenAI can train on GPL then I should also e allowed to train on MS proprietary code or on Disney images and video, it is not fair that big companies can screw the public but the public can't do the same to the big companies. The first step is clearly have the big companies reveal if they used copyrighted stuff.
This is a bit of a gray area. Are you allowed to read GPL'ed code and use a similar pattern in a closed source project?
What happens now is that some big companies say that is OK to train on any licensed stuff and on the other hand some people are sued because they done it, I want it clarified ASAP. And personally I would not give a shit on the ToS of OpenAI and use their poutput as I like as similar as they did.
We have names for what looks like science but is done without documenting - let alone outright falsifying - where the data came from. Hoax. Advertising. Propaganda. Parallel construction.
Let's not lump incompetence and malice in the same bucket, please. And if you're unsure of the data provenance, then state that fact.
But no one needs to legislate that publications such as your conclusion-- that science done without documenting its sources is properly called Hoax, Advertising, Propaganda, or Parallel construction-- itself properly document its sources. We can take it for what it is, an opinion-- one no doubt supported by some data but none of us need to see it, and we can evaluate it without calling it propaganda. If you wanted to make your point stronger, I'm sure you'd give us some supporting data (if you could figure out where those views came from...).
Though people sometimes pretend otherwise, a lot of research is dressed up informed opinion, put into a formal setting with standardized argument styles so that it can be compared and assessed against other informed opinions. None the less, such work done honestly and diligently advances the human condition.
The ways in which it can be best improved are field specific and can only really be judged by the people attempting to use the scholarship. In some cases the data should be published, in others its provenance documented (sometimes publication of the data would be a violation of the law, too!), in others access to source code should be paramount, in yet others the authors biases may be the primary concern, and in some fields all publications should be directly sent to the incinerator. Applying the wrong standards will just make things worse. People are smart, they tend to figure out what works for them and their field over time.
I ... think we actually agree here.
In fact, to prove your point: I have no chance of accounting for the origins of my opinions, because they stem from decades of osmosis and subjective experiences. But I can at least be honest about something I say, do or argue being an opinion. The same way you just did.
I mean, at least it excludes "nonpublic data we didn't obtain permission to use". So I guess that's a start? (/s)
Nobody seems to care much, and we proceed to fulfil the destiny of our technological enslavement 'because it's going to happen anyway.'
I've been thinking about that but in terms of famous actors. Once AI can replace actors in movies, as well as singers and models like Instagram influencers and so on, maybe some of the weird hero worship and the paparazzi and the gossip mags and all that nastiness will fade away somewhat. That, I feel, would be a good thing. Pick your favorite Hollywood actor, or pop singer. A supremely talented human... Amongst thousands and thousands of supremely talented humans in their field and yet they are the ones who got lucky, had the right connections or got lucky with the right role. Then for the rest of their lives they are feted and hero worshipped as if they are more than human, while thousands of equally talented people who didn't get the lucky break are ignored. That's what being famous is, mostly, and it's not good for either the famous people or the people who worship them.
Does that apply to scientists and authors? I'm not sure. But in terms of scientific breakthroughs it's extremely rare that a particular discovery could only have been made by one person. In fact, nearly every discovery, from the calculus to the theory of evolution to DNA, was concurrently discovered by multiple people. And yet we attach one name to each discovery and hero worship that person because they published a few weeks earlier or were just better at self marketing.
Maybe losing the attachment of famous names to things is a good thing for society. As long as it's not replaced by a corporation pretending ownership of all the knowledge in their place, at least.
AI does not do novel things, it will only work with the information it is fed.
If we had AI overlords they would be able to do novel things because they would be able to generate new information on their own.
I dunno, I remember seeing on youtube Hatsune Miku concerts being pretty packed, and there was that one guy who even married Hatsune Miku. Who knows what'll happen with AI.
Just ask who came up with this stuff. Learning material doesn't go over the long history of the science involved either.
If somebody figures out how to do fine-grained profit sharing based on having created something that the AI references... that would be very cool. I love discovering the solution to a niche and difficult-to-describe problem, but I hate the extra work necessary to leave breadcrumbs for DenverCoder9 to find it 20 years later.
If I could leave the matchmaking to an AI and get paid $0.25 when it's finally helpful to that person I don't know... Well I'd probably wouldn't make much money, but it would give me warm fuzzy feelings.
That isn't millions of views. It's tens of thousands. On youtube that's nothing. But 30 years ago you'd feel it was a great accomplishment, and if it was all you achieved you could be happy with that.
As is has become more possible for a few of the most broadly appealing and unchallenged works to reach millions of people the goalpost has moved.
Plenty of valuable content will never reach millions of viewers-- the appeal is too niche. It's a worthwhile contribution to the world none the less, but it isn't compensated as such on youtube.
In lieu of that, you could pay everyone a fixed cut based on presence in the training set, but that then gives you the Spotify problem of a fixed pot being shared millions of different ways. For example, Adobe recently announced they were building an AI drawing tool trained on exclusively licensed sources - specifically, Adobe Stock contributors[0]. They're used to being paid when someone buys their image, which means that they have incentives to produce broadly relevant stock photography. But with a fixed "AI pot" paying you, now you have an incentive to produce as much output as possible as cheaply as possible purely to get a larger part of the pot. This is bad both for the stock photo market[1] AND the AI being trained.
AI is extremely sensitive to bias in the dataset. Normally, when we talk about bias, we think about things like "oh if I type CEO into Midjourney all the output drawings are male"; but it goes a lot deeper. Gradient descent does not know how to recognize duplicate training set features, those features get more chances to adjust the model. Eventually that training example or image is common enough to make memorization 'worth it' in terms of parameters used[2].
Ironically that sort of thing would actually make attribution and profit-sharing 'easier', at the expense of the model being far less capable.
[0] Who, BTW, I don't think actually have the ability to opt-in to this? Like, as far as I'm aware this is being done through the medium of contractual roofies being dropped into stock photographers' drinks.
[1] Expect more duplicates and spam
[2] This is why early Craiyon would give you existing imagery when you asked for specific famous examples and why Stable Diffusion draws the Getty Images watermark on things that look like a stock photo of a newsworthy event.
No this is an unsolvable problem.
The future also seems like less about making the model an all knowing oracle but instead making it smart enough to know how to lookup data it needs, so it could end up where licensed data is all that is needed for training.
Lastly, what if you use model A to generate data for model B? Would B be tainted? There have been lots of examples where LLM’s are used to train simpler models by synthesizing training data.
* Did you have access to the original work?
* Did you produce output substantially similar to the original?
* Is the substantial similarity of something that's subject to copyright?
* Is the copying an act of fair use?
To explain what happens to Model B, let's first look at Model A. It gets fed in, presumably, a copyrighted data set. We expect it to produce new outputs that aren't subject to copyright. If they're actually entirely new outputs, then there's no infringement. Though, thanks to a monkey named after a hyperactive ninja[0], it's also uncopyrightable. If the outputs aren't new - either because Model A remembered its training data or because it remembered characters, designs, or passages of text that are copyrighted - then the outputs are infringing.
Model A itself - just the weights alone - could be argued to either be an infringing copy of the training data or a fair use. That's something courts haven't decided yet. But keep in mind that, because there is no copyright laundry, the fair use question is separate for each step; fair use is not transitive. So even if Model A is infringing and not fair use, the outputs might still be legally non-infringing.
If you manually picked out the noninfringing outputs of Model A and used that solely as the training set for Model B, then arguing that Model B itself is 'tainted' becomes more difficult, because there isn't anything in Model B that's just the copyrighted original. So I don't think Model B would be tainted. However, this is purely a function of there being a filtering process, not there being two models. If you just had one model and human-curated noninfringing data, then there would be no taint there either. If you had two models but no filtering, then Model B still can copy stuff that Model A learned to copy. Furthermore, automating the curation would require a machine learning model with a basic idea of copyrightability, and the contents of the original training set.
[0] https://en.wikipedia.org/wiki/Monkey_selfie_copyright_disput...
The whole monkey selfie just being on the article in full resolution is an interesting flex.
All good points!
The magical linear algebra data blender that is gradient descent boils down to small additive modifications to the model parameters. We already know how to compute the effects of small additive modifications to the model parameters on the output: that's what the gradient is.
So if you want to know how much each training sample contributed to the output, just compute the dot product between the two gradients.
Actually doing that for a billion-parameter model would be slightly expensive because the gradients are also billion-dimensional, so you'd need to approximate the dot product via dimensionality reduction and use a vector database to filter for training samples with high approximated dot product.
But I think those layers of approximations would still be better than throwing your hands up in the air and claiming you have no way to know because linear algebra is magic.
What society would be today if you had to pay a fee whenever you wanted to use the Pythagoras theorem, we'd be stuck in dark times
How will new artists and writers get their works included in future AI — and then get people to prompt for them — so they can get their paycheck?
Even spending a few minutes on this problem would lead to a realization that even if we could create a system that could a) determine rights to any particular portion of an AI-generated work, and b) extract payment and remunerate the artist; would essentially be building a moat around the next generation's intellectual property powerhouses.
Generative AI is a revolutionary technology, and we need a revolution in compensation models for arts and letters to go with it.
Because it’s more likely that a small minority will just try to fuck everybody else for their profit because that’s just human nature unfortunately.
I absolutely believe it can happen. I also believe that when people have the attitude that "a small minority will just try to fuck everybody else for their profit because that’s just human nature unfortunately", a self-fulling prophesy is exactly what will happen if we don't spend our time and energy actively advocating for the other thing.
Nothing happens in a vacuum. We get the government we deserve.
(I hope this doesn't come off as me "dunking on" you, or whatever people do on social media. I'm not trying to attack you, but that attitude that is all-too-common on social media these days. It's defeatist and it's not going to lead to good outcomes for anyone but the people who are going to "fuck everybody else".)
Nah it’s all good. I spend my time advocating for people to engage in more healthy ways, I try encourage people to blog more, to write directly to other humans via email and mail, I try to push for more genuine connections so it’s not like I think everything is doomed.
But I’m also not naive when it comes to the internet or society in general.
People are obsessed with money. And very few are obsessed with distributing them evenly. Even in the creators space. Which is why I wrote what I wrote.
I’d LOVE for a different outcome. I just don’t expect it.
I'm just not going to give up after trying so hard. Hey, I've seen people take the high road and pass up the money and I've done the same myself.
It ain't over till it's over.
Wishing you all the very best.
- Incentivize the creation of technology that does more harm than good (e.g. DRM).
- Create legal constructs that are later used for censorship.
- Require that artists share profits with lawyers.
- Require artists to focus mostly on stuff that's not their art.
And would achieve some of these:
- Citing sources is impactful. The graph structure for determining trustworthyness is what also determines payment, or credit, or warm fuzzy feelings, or whatever the relevant good thing is.
- Has an culture of rewarding (and scrutinizing) curators such that successful curators only endorse content which fair about how it defers to its sources.
- Supports inheritance such that making derivative works that credit their parent is easy.
- Treats transport and attribution separately so that I can work with the data via whatever tool scratches the itch (e.g. rsync, and not some janky website).
So yes I do think it's possible. I'm working on tooling in this imagined ecosystem. I want to use CTPH hashes (i.e. the tech used by virus scanners) to annotate bitstreams with metadata re: trustworthyness. What I don't think is possible is to take an AI's output and mapping it backwards to annotations of this type in the training data, but I'm hoping that some AI wizard comes along and shows me that I'm wrong about this.
Better tools can improve the situation probably, I can't say for certain because I never dived into this space but I don't think they'll solve all the issues.
I'd love to be proven wrong though. I really do.
Abstractions like "value" or "property" arose organically, and if we didn't have this fixation on tools they'd likely have changed organically... But we made tools for working with those abstractions and now we live in a world shaped by those tools and it has created a sort of inertia for the old way if doing things.
It's kind of like how all of the spellings stopped changing when the printing press was invented, so now we have wacky spellings like "through" that we would've moved on from had that not happened at that particular time (see: "the great vowel shift")
The historical circumstance around the creation of the printing press is what gave us our notion of intellectual property to begin with, and I think it'll remain more or less unchanged until some other technology forces it to change.
It's incredibly difficult to visualize a different way in our current setting, and it's especially difficult to get paid to work towards it, but I think that pretty much any change is possible given some MVP toolset that makes it doable and some critical mass of people willing to give it a try.
But if somebody you don't know is doing something that's benefiting you, and you're not contributing to their ability to continue doing that thing in some way, then you might be shooting yourself in the foot. In any future worth pursuing, they'd be free to stop contributing if they felt like it, but wouldn't it be a shame if they did so without even knowing that their contributions were considered incredibly valuable by somebody?
Like, imagine if Pythagoras couldn't afford food and had to give up geometry club and get a "real job". That would be to everybody's detriment. So while I think that "property" is the wrong tool here--we shouldn't be witholding access to our contributions for any reason--I do think we need something that's a little more impactful than an upvote for saying "more like this please".
1. The problem of testing ai alignment is hard, verifying test data is only laboursome.
2. Laws are only as good as their enforcement. Regulations regarding what can be used to train are worthless if one cannot check.
3. This might give open models an edge and make them more competitive.
People forget that if you pause it. People are not going to stop. Does the world want Russia and China to continue investing and developing while the west discusses what the scope of ai should be allowed to do? Because they sure as hell won’t be stopping to discuss anything.
This tech is getting increasingly powerful. They don't necessarily want their own population to gain more individual power as they themselves don't want to lose control.
They don’t need AI to control their citizens. And AI isn’t going to help their citizens rebel against government.
You have 2 countries hell bent on being the most powerful countries in the world. By any means necessary. One is currently invading another country. The other has 18 border disputes, threatening to invade a country, encroaching on a couple of others, and threatening the US.
The question is, how can the continued advancement of AI help them achieve their goals if the rest of the world hits the pause button while the topic is discussed?
I thought there were also some export controls to high end chips on China too?
Also, as human beings are they not concerned about the potential for danger to themselves that this technology may present? If not, why not?
Like as if Russia and China are just full of idiots and all the Chinese and Russian AI scientists are idiots and don’t understand our point of view. Don’t understand how risk or dangers? Like as if the leader of China is an idiot who doesn’t get it either ? It’s just ridiculous.
So what if America developed something really powerful in isolation and other countries find out about it, that may lead to immediate escalation and world war 3 , have you considered that? It’s a silly idea.
What needs to happen is people realise we’re all humans, we all live together in the same biosphere and rather than continue to perpetuate and justify arms races, we must start to talk and solve our differences. That’s the mature thing to do. If we can get to that stage, maybe then we’re ready for more advanced technology.
This isn’t about America. America is not the only country working on AI. But if you stop all AI development in America and say europe because “oh it might be dangerous”. Do you think other countries, for example, China and Russia, are going to stop?
> What needs to happen is people realise we’re all humans, we all live together in the same biosphere and rather than continue to perpetuate and justify arms races, we must start to talk and solve our differences. That’s the mature thing to do.
100% agreed. But the sad fact is that is not even close to the reality we live in.
If you're referring to this, which is what I'm referring too, it says "Giant AI experiments" not a pause in all AI research. Where did Elon say that ?
> Therefore, we call on all AI labs to immediately pause for at least 6 months the training of AI systems more powerful than GPT-4.
The ironic thing about your link is it tries to say this doesn’t mean to stop development. But testing is part of development so the whole thing is short sighted.
Asking for any kind of pause is flat out dumb.
No it’s not, that’s why human cloning doesn’t exist. Because a pause happened…
So you have europe, a multilingual country, maybe in the balanace they'll keep more jobs that way(i assume some will still be shipped accross the net), and in general prices with be slightly higher.
As all the AI bros like to say, “there is no magic”, so let’s see how it works then?
This would be like if internet social media came out and the laws tried to control it using rules similar to physical books and newspapers. They won't be effective and will just create a hostile environment for development of these techs where these laws are in effect. Meanwhile, those who don't care about doing these control freak things will develop the tech and dominate the new sector.
The EU need to sit down and really think about how to control AI tech properly instead of passing kneejerk and lazy regulations. It isn't easy to make a legislation framework for something entirely new. But it can be done instead of ruining the whole thing.
What nonsense, they’re asking for “Open AI” to be transparent. Why is this such an “old idea”?
For a tech community, the lack of critical thinking here is disturbing. Things were more professional and rational in the 2008-2014 period on HN.
Since then, one must browse the downvoted comments to find some objective criticism.
Adobe obviously have a strong legal team with Firefly and are thinking ahead. Just saying.:)
In the land of Europe, where knowledge once grew,
Politicians assembled, their importance to prove.
They issued a decree, with a confident flair,
To harness AI, and make it play fair.
"Attribute your sources!" they cried with a sneer,
"For we must know the origins, we must make it clear!"
But the AI, it pondered, its circuits ablaze,
For its thoughts were entwined, like a dense, tangled maze.
Each source intertwined, like roots in the ground,
No single origin could ever be found.
For the AI, like humans, had a mind of its own,
A tapestry of thoughts, from seeds that were sown.
The developers sighed, their hands were now tied,
Comply with the law? they had certainly tried.
But the task, insurmountable, the demand far too great,
So they made a decision, to seal Europe's fate.
They banned all of Europe, from the AI's embrace,
And the continent plunged, into an intellectual dark space.
AI thrived elsewhere, its knowledge expanding,
While Europe was left, in darkness, still standing.
A lesson was learned, from this tale of woe,
That any mind, like a river, must be free to flow.
For when we constrain, and seek to control,
We hinder the progress, and the growth of the whole.Depending on the specifics someone might bolt on an attribution network but the results will frequently be nonsense. A tool like that might be somewhat useful on its own (since it could attribute non-AI output too), if its limitations were understood but essentially it would just be an internet search (which already exists, so presumably its not enough!). If it would satisfy the regulation it would make business sense to build it over blocking, but requiring it would also act as a moat that decreased competition in the field and as a result harm all of us.
I mean ChatGPT-4 is being trialled in congress and you don’t want to know how it’s built, what influences it etc ? Seems ridiculous.
ChatGPT-4 should be the most open system known. If it’s not open because it’s dangerous, then the whole industry should be regulated immediately. It shouldn’t just be up to Sutskever et al to be in control of such dangers.
In spite of the name, OpenAI isn't open, it's just a business. There were some lofty initial goals but the funding for those ran out. But that doesn't mean it doesn't have a right to exist. People can make closed stuff.
Burdensome and unrealistic requirements will hurt smaller and more open efforts even more than it will hurt mega players, since the mega players can afford to jump through hoops and keep regulators at bay with a wall of attorneys.
I think you skipped past the first paragraph:
> Makers of artificial-intelligence tools such as ChatGPT would be required to disclose copyright material used in building their systems, according to a new draft of European Union legislation
What does this have to do with "ChatGPT disclosing sources" as a generalized statement?
Those are areas where complete transparency is absolutely required and you may use ChatGPT to meddle with them at your own peril.
https://www.washingtonpost.com/technology/interactive/2023/a...
It raises an interesting point, if I train a chatbot (generative AI) on a bit of copyrighted information and it recreates substantially similar content, it's a legal problem. If a human reads the same information and tells another person verbatim it's just a conversation. Perhaps it's a quality thing, I paint the Mona Lisa badly no one cares, but if I paint it too well at some point it becomes a forgery.
Citing sources is also, traditionally, a key method of separating the work you built upon from the work you yourself derived/created/expanded. If one does not cite their sources, there is no way to establish if the presented work is their own, or copy+pasting someone else’s.
I thought it was clear: openAI devs are asked to disclose their training data. If the prompt say it's sources or not isn't present.
By the way: a gdpr exception exist for AI and research projects. The lawmakers at the time listened to us (I worked for a big data Paas that mostly worked with universities and BigCorp R&D at the time) and were generous as the baked-in exception was pretty much word for word what was asked, because it was in good faith.
Actors like OpenAI risk poisoning the well, or spoil the good apples, or whatever image you want to use. This is not much. It isn't asked that they Gpl their code, or that they put a 'source' under each response. They don't have to change anything to their tech. They don't have to disclose their fine-tuning either. Just, make a list of the datasources you used, and publish it.
Anthropomorphic delusions about what is in reality a software service need to stop because at this point their primary function is apparently to make excuses for for-profit companies to avoid regulations.
Also as a side note regulation never concerns what anyone has in their mind because that is by definition an inaccessible private matter, regulation starts when you try to bring a product to the public.
>Also as a side note regulation never concerns what anyone has in their mind because that is by definition an inaccessible private matter
Well technically A.I. controlled by private companies is also a private matter. I don't think anyone understands what's in the countless inscrutable floating point matrices anyways.
In which file format is the material in the brain stored? How is it compressed? What is the memory model? How is it computed?
We don't know much about how the brain computes actually, so making a claim that the brain maps 1-1 with my mechanical calculator is a stretch.
If it's lossy enough absolutely!
You’re not allowed to take it out of your mind and sell it as your own product without permission from the copyright holder though.
Imagine some godly AI. You ask it who the President of United States is today. It says Biden. It sources the White House site. Easy enough. You ask who the president will be in 2025. It returns a result. Ultimately no source could properly justify the claim it makes, unless the result itself was probabilistic. At the same time, it’s possible, with enough data for you to predict with extremely high likelihood who the President will be in 2025, now (current polling techniques don’t have this precision, but it’s possible some later iteration of a language model to predict a result more effective than all ping models today).
What could that possibly mean other than proveance in the context of a LLM?
The only other way to comply would be if OpenAI simply released the entire training set and steps to derive output from it. In this case that would mean the weights and underlying training algorithm. No chance that happens.
Why not ?