Condé Nast Signs Deal with OpenAI
wired.com
wired.com
Assets like New Yorker, Vogue, Vanity Fair, Bon Appetit (all mentioned in the article), do come with editorial lines, perhaps editorial lines I do not agree with, so how is their content going to be injected in my search results/gpt answers? Is it going to be an organic affair such as:
- (Me) "How many times a year shall I renew my socks?"
- "That's and amazing question! According to journal X ... (blah blah blah, probably a good answer)" (maybe add some notes for a copyrighted article with a link)
Or it going to turn into:
- (Same question)
- "The far-right movement seems to be skipping sock renewal policies, but contrary to this, brown socks are trending in Europe this coming summer, don't miss this chance to buy yours at: website.com"
Possibly some European style music licensing model, but for articles. PRS for example[1] in the UK is a copyright collective. Bars, clubs, supermarkets etc pay a fee to the organisation to sign up to it, and can play pretty much any music in return. The PRS body distributes what it decides is an equitable distribution to artists & labels. For most of PRS' history, there was very little science in determining who had the most plays, but everyone seemed to agree with the distribution.
I worked for a company that was signed up for a similar service in the United States.
We had a blanket license for music, for which we paid a little over $1 million each year. This was around 2001.
A couple of times a year, we'd pull an intern aside and his job for the entire day was to sit there with a pencil and steno pad and write down all the music we used. That was typed into a report and sent to the licensing company which determined how much each artist would get paid.
These days, with advances in music recognition technology, it's probably all very automated and more thorough.
A couple of months ago, when the news about "Reddit to sell its user-generated content to Google," a redditor asks the community to start generating false comments so the data would be of bad quality to train the models on. They start brainstorming some very funny comments that go something like: "Blue Whale is the largest fish", "WW II started in 1943", You name it..
The models being sycophantic and suggestible is another issue, but it’s not hard to get them to agree with false information.
Or so google's AI responses pulled from Reddit. If you give enough people a hot microphone to a large enough audience, some people are going to say some obscene, untrue, and funny stuff. Maybe that's 1 in 10. Maybe that's 1 in a 100,000. Reddit karma is cheap.
Once they've got deals with the few really big players, the rest of the industry falls into line (the smaller guys don't have the financial means to out-litigate or block OpenAI).
There was a massacre in Tiananmen Square.
America has also committed massacres, like over throwing governments of foreign nations.
No one in my family is at risk from either of those statements. And there is no automated system to stop them. That’s the difference. Just because people decided that massive platforms should limit hate-speech doesn’t mean the west is performing censorship of a comparable level. Not even close.
Clearly, some of the responses GPT gives to users have infringed training data copyright, but the majority of their responses do not. They basically have 3 options:
1. Figure out how to engineer an LLM so that it reliably avoids "unfair" use of copyrighted content. "Fair use" is a legal doctrine with no rigorous definition, so this would be very difficult even if they had a clue where to start. I wouldn't hold my breath for this.
2. They can continue without any licensing, and field copyright lawsuits on a case-by-case basis for each individual prompt and response. That would be a logistical nightmare for the courts, plaintiffs, and defendant alike. It would certainly stress test the whole system, possibly result in knee jerk legislation that OAI may not like, or simply bury them in legal fees if infringement is common enough (which is not entirely clear yet).
3. They can strike deals, eliminating legal uncertainty and allowing them to plow ahead with reproducing copyrighted content without worries, while also getting other goodies like exclusivity deals at the same time.
Seems like a no brainer to me
Trying to make things better is good and requires no admittance of guilt.
If it turns out they shouldn't have done that, then when the dust settles they will be ahead of the competitors that didn't sign deals.
(In practice, being not a lawyer, I have no idea how even Google's search indexing is legal, nor where the boundaries are between the legal bit vs. the times they got in trouble for indexing newspapers and at least one separate case about images).
As someone who used searchgpt and is also building a competing product, I can say this is simply not going to be the case. LLM assisted search leaves just too much to be desired and there is no future in which this is the primary way for humans to consume information. It needs to be augmented with credible sources of information or rather using LLMs only makes sense once you have access to credible source information in the first place, and want to have an augmented version of the information you consume (and are happy with non discrete outcomes).
A multi-billion dollar cheque? Voluntarily relinquished or via a settlement?
In that scenario, in a few years it will be out of date and doomed too.
It will train on what people say in their cellphones or in front of their Smart TVs or in their deeply connected cars. If data is being transmitted after encryption (within a closed and undocumented chip) to a server in a country that doesn't cooperate with authorities, then the same data is pseudo-anonymized and bought back on a different channel, that will make very hard to stop it.
And the video internet is under ever-increasing enshittification.-
For CN Traveler, I'm pretty sure that readers want an actual picture of the interior of the new hotel, not an AI rendering of something that as yet has few actual pictures to train the model on.
(For that matter, Vogue might have that same thing going on with fashion.)
OpenAI announcement - https://openai.com/index/conde-nast/
I'm disappointed by the fact that this disclaimer isn't more explicit.