Associated Press clarifies standards around generative AI
niemanlab.org
niemanlab.org
Why are they paying or even asking for permission for training on data?
DALL-E helpfully slaps a ShutterStock logo on some creations, for example. https://twitter.com/amoebadesign/status/1534542037814591490
Seems like you could train the AI to recognize the logos and edit them out?
No it isn't lol. a shutterstock logo is just more common ground for the model because of how often it will appear in the dataset.
They would say they are owed licensing fees regardless of whether the shutterstock logo appears in the output. The appearance of their logo in the output merely proves that the system's outputs are derived from their copyrighted images.
But upon the absence of the mark, shutterstock would argue that the provenance of the image WAS shutterstock?
I don’t think this is the salient feature of this phenomenon, but the proximity of these two arguments is interesting.
The internet has shaken the system but the content producers were able to adapt, albeit resulting in lower quality content.
However with the raise of AI, the thing completely shattered. Previously, someone reading the content and telling it to others wasn't a problem that breaks the compensation scheme for content producers but with ChatGPT and similar we have a situation where this "person" can tall about it to literally everyone. Some new compensation scheme is needed and OpenAI is probably trying o act as the "nice guy" to prevent the urgent need of a scheme that might limit their ability to consume other people's content.
This stuff is a legal minefield, as a for profit company building their core product it’s very difficult to argue each of these is fair use etc. Though I am sure that argument will be made, it’s risky when dealing with companies whose businesses model centers around IP.
There's a reasonable strong argument that crawling public pages for "indexing" (aka learning) is fair use based on the Google precedent and case law from ther early 2000s.
The argument is much less strong if those records aren't available.
There are good reasons to keep confidential info out of LLMs that you don't control, but I'd think it would make sense for anyone to run text through a locally-hosted LLM for editing suggestions and the like.
But I can see the potential for harm in over-humanizing by a news outlet. I'd be interested to hear what their decision-making process was for this point, if it was obvious or if they went back and forth, what arguments they had for which direction, etc.
Really? I would have said the opposite - journalists have no obligation to parrot what companies' marketing departments feed them, and in fact usually ought to do the opposite.
Russia might name its invasion of Ukraine "Anti-Nazi Operation Freedom Eagle" but you wouldn't expect a war correspondent to repeat such obvious propaganda. In general, journalists have no obligation to follow companies' and governments' naming preferences.
Or LEGO gets written in all caps.
To me, the style guide decision to conform to the desired persona when describing it is along these lines.
An LLM hopes to be an "intelligence" that can understand and manipulate text along the logical boundaries of language; and do so intentionally.
What an LLM really is, is an inference model that can reorganize text across boundaries that closely "align to" real language patterns. This is accomplished by creating a completely new pattern (the model) from inferring whatever patterns already exist in the training corpus' text.
A Large Language Model (LLM) serves as an alternative to true language comprehension. It is not an equivalent replacement, nor does it intelligently navigate itself with any explicit intent.
The act of "intelligently navigating the content of language" is at the core of journalism. It's incredibly important for journalist to both recognize and articulate the difference between an Artificial Intelligence realized, and any technology in the category of AI pursuit.