Major Sites Are Saying No to Apple's AI Scraping
wired.com
wired.com
I’d argue the way Wired is reporting this as a gotcha is also a conflict of interest because of their deal with an effective competitor.
https://www.wired.com/story/conde-nast-openai-deal/
> “WIRED can confirm that Facebook, Instagram, Craigslist, Tumblr, The New York Times, The Financial Times, The Atlantic, Vox Media, the USA Today network, and WIRED’s parent company, Condé Nast, are among the many organizations opting to exclude their data from Apple’s AI training”
Versus
> “The media company joins The Atlantic, Axel Springer, Vox Media, and a host of other publishers who have partnered with OpenAI.”
I have my own serious ethical concerns about AI, but from a pure searchability standpoint it is likely to cause more problems than solve.
Information is going to start being silo'd. It would be like if I had to go to Google for one thing and Bing for something else.
That being said, I am also quite concerned about what you mentioned. I was not happy when I found out that ArsTechnica signed a deal with OpenAI (or I assume their parent company did). Did that deal include any requirements of how they talk about OpenAI, AI in General, or their competitors?
The concerns over AI are real, but they are putting a short term gain over the long term damage by somehow thinking that licensing data is going to save them.
The tone of this is a bit weird, do we have similar "blocking google", "blocking anthropic", "blocking meta" articles? Instead of what is really happening is those partnerships were formed so everything not that partnership is being blocked.
AI just blows all that up by providing a general-purpose plagiarism engine. It offers a substitute for any content, with a similar-looking but inauthentic version. People might accept not paywalling in order to reach human users, but nobody wants to be fodder for the plagiarism bot.
In some ways this is the "Can Googlebot scrape your content and display it on the home page but without generating any ad views for you?" debate we've seen with news publishers before.
Personally the way I see it, you can't argue about being worried about plagiarism while allowing one company to just mine all of your work.
You either block all or you allow all. Anything else is not helping either situation.
Ars since you mentioned it https://arstechnica.com/information-technology/2024/08/opena...
To your point about siloed information, we also see this with Reddit and Google, since they have an exclusivity agreement.
There is an argument for banning (or limiting the term of) exclusive licensing arrangements for purposes of training an LLM. Nobody is forced to do business with anyone. But OpenAI or Facebook shouldn't be able to force a publisher to only do business with them, thereby closing the market to new entrants.
we're already there: a month ago reddit signed an agreement with Google to ingest their data, and then blocked everyone else (e.g. bing)
https://www.reuters.com/technology/reddit-ai-content-licensi...
AI killed the Internet
a company sends a robot to watch it. 7 dollars for a single ticket.
the world spends their ticket money on hanging out with the robot to hear about the movie.
i can't make heads or tails of what's next. i am as fascinated as i am petrified.
If the laws are written in code (pun intended), they would look like:
if typeof(subject) == Human then
...And if we want to jog down this rabbit hole, I suggest starting with existing animals which we're realizing are sentient, but still prevent them from going into movie theaters. Once octopi can visit the movie theater, let's revisit the the robot argument.
That's not why. We don't know what test could be performed to determine if some future (or present) AI is or isn't sentient, we just assert it and get on with our lives.
No, the reason robots and software have no inherent rights is because the law says so.
As a demonstration proof that this is a purely legal status that has nothing to do with underlying nature, the co-inventors of the automobile did not have many rights afforded to her husband: https://en.wikipedia.org/wiki/Bertha_Benz
And conversely many such rights do exist for corporations, which are not themselves sentient.
This is based in a limited theory of rights being delivered by a state as sovereign rather than inherent to every being. There are other valid ways of interpreting rights beyond "because the law say so."
I'm not saying you're wrong, those kinds of debate are great at telling you what the law should be, but the practicality of it is that a right you can't enforce is not useful.
Influence is power. For a stark example, consider Sharia.
You're right that using sapience is a bad way to identify what is and is not a species though.
This status difference is rooted in significant recognized differences and philosophical beliefs about the value of individuals, and in the idea that the law and society should exist to support that. It's pretty far from "purely legal" and the fact that we're going to let people own robots should be only one of the many clues that there's something different about them.
All that philosophical stuff, that only mattered because it resulted in a change in the law — and in the case of the US different philosophy led to a civil war to reject/enforce the law.
Wait, whats the threshold for sentience and do all humans meet that standard? Does the metric for the threshold depend on cooperation from the subject? Asking due to recent headline "One-quarter of unresponsive people with brain injuries are conscious." [0]
It seems more straightforward to make the speciesist argument until cyborgs start containing majority biological components.
A service dog can be brought into a theater just fine, as it is sufficiently intelligent to be trained to behave appropriately.
There's nothing that specifically prevents animals from entering a movie theater, I'ma say they mostly just generally lack 1) any intellectual desire/intent to watch a movie, 2) appropriate social behavior to cohabit a space with humans, 3) money to pay their ticket.
---
I would agree that current ML/NN based models might seem like a way to more or less circumvent copyright, but don't feel like there should be a blanket ban discriminating against potential intelligent entities that might arise in the future from more sophisticated technologies.
I also wasn't completely serious with my initial post. It is rather enjoyable to take an idea to the extreme and see what people's thoughts on it are.
>argues there’s no evidence
But for the moment, and I say this as someone very impressed and somewhat scared with AI progress, they're probably sufficiently un-person-like today that it's not unreasonable to ban current models.
Ants may not be taken seriously if they were to offer to pay, but then again nobody will stop them just walking in to a theatre without a ticket either.
I think the dystopian future we're heading towards is personalized phishing and/or scams that sends to your hijacked accounts' contacts that you're going through a hard time and requesting donations, using training data from crowdfunding sites.
Or, one-level more dystopian, hijacking social media accounts and advertising AI-generated Patreon-style content using the actual account owner's likeness.
What's next is the bot writing the whole story and generating the visuals. For the cost of one actual movie, Slopflix could turn out thousands of recap style stories.
Yes, but I can already do that by reading plot summaries on Wikipedia.
I've read precisely zero Harry Potter books, and of the films seen 1, 2 in the background muted, 3, and 4.
Despite this I know about Dolores Umbridge's personality (oh, and just recognised the reason for the name), that Snape Kills Dumbledore, and several other events and characters that were also not in any of those films.
(Alternatively: https://imgflip.com/i/91shhp)
I am also thinking (more and more lately) the movie The Congress (https://www.imdb.com/title/tt1821641/) and how easy it will be very soon to make a company without human actors. The mega-big studios can train AI by feeding it all their movies, then feed it with the scenario (this is where we - the humans come in - the famous "prompt engineers"), and a movie will be created.
Now, to the above scenario there will be a lot of corrections (can't walk in the air/water/nails)(can't open with killing a dog unless "John Wick" or "Sacred Games")(Sandra Bullock can levitate in "Gravity" but not on "Speed")(etc.)
Eventually if you hire 1000 people to watch 2 iterations of the new movie per day and make corrections to each (100025 = ), the AI will learn and improve on these unrealistic outputs and eventually you can produce movies weekly.
Yo Netflix people, if you haven't started doing this already, START! :) (oh and make a follow-up series on Sacred Games while at it)
EDIT: that's quite a tangent, but I can imagine the day where I will be asking my Paramount-AI "hey, make an original Star Trek series based on this-and-that, I want to binge it next weekend!" (or show me the one you created for someone else)
Still wondering where AI is gonna get its raw information from in a few years when 90% of websites have given up updating.
Facts, opinions, or skills?
Facts only need a single source of truth, skills can be learned in sim.
Plenty of room for a gossip column that keeps robots out, but how much will that matter?
I do agree with the point in a different way however, as spamming a fake reality is quite capable of undermining the world models of us humans.
Two modes of censorship, banning true things vs. overwhelming the signal with noise.
Gossip is a noise we're very interested in.
For those who make money on subscriptions it is a no issue, since the robot can't index the content anyway.
But yeah, if you are an influencer, well the world is probably strictly better of without you.
If AI ever becomes a major way to search for businesses it will quickly become pay to play. Letting an AI company scrape your website won't get you much at that point beyond the ability to pay for traffic.
At least search engine crawlers may drive traffic to them later. What’s in it for them to contribute to model training?
So basically just google? I don't think other search engines matter that much. I've never heard people specifically doing SEO for anything that's not google.
If they just use my content to then tell users without needing to come to my site, and get ad revenue from that - no thank you. Not what AI is doing now, but looking at SEO I don't see where else this could be heading...
They are saying Yes to everyone. There is no indication anywhere that they block AI scrapers.
In fact, they let a blogger 'profit'.
https://wordpress.com/blog/2024/07/30/perplexity-partnership...
They just don’t want the competition.
It'd be much more concerning if academic publishers, code repositories and wikipedia refused scraping.
Apple drank the kool-aid and now has to deal with the fact that those bad faith arguments make no sense in reality.
This is a meaningless semantic punt to the word public. At the end of the day, there is a real debate around how ownership works in relation to AI.
I don't know how long it's going to work like that, internet is basically the farewest. If you can't paywall your data, someone is free to take it.
Beside I bet most article written on these website used artificial intelligence to do so.
Do you think there's something wrong with having a ToS that specifically excludes AI?
Content on websites has copyright in some form, just because it isn't behind a paywall/accountwall doesn't make it public domain.