Reddit will begin charging for access to its API
nytimes.com
nytimes.com
> Reddit’s API will remain free to developers who want to build apps and bots that help people to use Reddit, as well as to researchers who wish to study Reddit for strictly academic or noncommercial purposes.
> But companies that “crawl” Reddit for data and “don’t return any of that value” to users will have to pay up,” Reddit co-founder and CEO Steve Huffman told The Times.
But they do return value to users. I'd much rather get my answer from a Chat-GPT query than scouring through Reddit.
Maybe he meant that they're not returning value to Reddit in which case he'd be right, but I hate him trying to spin this for the users.
Notice that 'to users' was outside of the quote. That was an editorial addition.
> “Crawling Reddit, generating value and not returning any of that value to our users is something we have a problem with,” Mr. Huffman said. “It’s a good time for us to tighten things up.”
So it is "to users" but more specifically it's to our users.
I would agree with Huffman here: crawling the data to build ChatGPT gives the value to ChatGPT users who aren't necessarily Reddit users, and by short-circuiting queries and processes that otherwise may have led to new Reddit users, it's taking value from all Reddit users.
So the question is: how significant is the union of those two sets?
Just because you want to use Reddit data, doesn't make it a Reddit user, does that make sense?
It's not possible to hide crawling at a large enough scale, right? At some point, certain IPs/user agents will (should be?) hit with CAPTCHAs to be able to have access to content and no amount of user agent/cookie/session/whatever spoofing will get around that, yeah?
Against a less-savory actor using hundreds of IPs from residential proxies/compromised hosts, you’re gonna have a rough time, especially if you’re unwilling or unable to use aggresive fingerprinting or (vomit) CloudFlare. Not to mention CAPTCHAs are generally already a solved problem for scrapers.
You could set a minimum karma threshold, but that would only promote karma farming; which is already widespread.
I wonder what their monthly active users look like if you filter out 1 person switching through 3 usernames/accounts for example.
https://news.ycombinator.com/item?id=32912075 https://news.ycombinator.com/item?id=17750801 https://news.ycombinator.com/item?id=22109969 https://news.ycombinator.com/item?id=30764757 https://news.ycombinator.com/item?id=29839960 https://news.ycombinator.com/item?id=22406277 https://news.ycombinator.com/item?id=23897705 https://news.ycombinator.com/item?id=34639212
(Hypocrisy disclaimer: I have sites behind CloudFlare.)
As one example, I can no longer browse the site for Lowe's (big box home improvement chain). Consequently, I now buy everything from Home Depot (their competitor).
It's astonishing how Cloudflare can do such a poor job of determining the difference between a potential customer and an attacker. Life's too short to solve captchas for an intermediary, so I don't bother, I just find a competitor who wants my business.
I don’t find that astonishing at all. I can’t see how you’d disambiguate someone who is anonymous for good versus bad reasons. Not supporting the death of the anonymous internet, but it’s not happening because of incompetence.
My outsider's impression is that Cloudflare has decided to rely much more heavily on browser fingerprinting than on classifying good/bad network activity. That puts them at odds with anyone that's taken steps to oppose being monetized by advertising firms.
One obvious clue would be that there are no attacks coming from my IP address.
Put another way, one could reason that they'd prefer to do business with Lowes because they are actively investing in security measures. Perhaps your data is more likely to be compromised at Home Depot.
For companies that make money by having more MAUs, well, yeah, they're going to have a real "rough time" detecting inauthentic traffic
The most “reputable” company in this space is Bright Data (formerly Luminati).
BrightData is the biggest of them, they run the free VPN Hola, and have an SDK app owners can install in their apps that allow selling bandwidth from installs. For someone who is price sensitive, trading some free residential bandwidth for whatever service is pretty compelling.
I'm sure there are scummy ones, but Bright seems to require pretty explicit consent. Not affiliated, just looked into it for some apps I have, but the payouts weren't good and I didn't think it'd be a good fit for our users.
Captchas are super easy! There's a gagillion captcha bypass services for every type of captcha. Just snag the captcha token, send it in an API call, and then you get a verified captcha token.
See CGNAT for more details about mobile networks. https://en.wikipedia.org/wiki/Carrier-grade_NAT#cite_note-of...
It's pretty much impossible to stop the top 1% of the most dedicated scrapers without affecting end user experience.
Only if the connection is over IPv4.
The mobile networks were among the first major adopters of IPv6, and most now give each device a unique IPv6 address.
tl;dr Dedicated crawlers built by sophisticated actors are more or less impossible to defeat.
At this point computers are probably already better at solving them than humans are.
Sort of, generally, except it's a lot more complicated
It's complicated is the short answer.
(Edit: merging https://news.ycombinator.com/item?id=35618695 hither now)
Depends on if the changes affect pushshift's crawling.
I'd be fine if they didn't. The ratio of useful bots to annoying ones is very low.
So yeah, Reddit may not need bots, but refusing to allow apps is just pushing another nail in the coffin of competition.
https://old.reddit.com/r/apolloapp/comments/12ram0f/had_a_fe...
> There was a quote in an article about how these changes would not affect Reddit apps, that was meant in reference to “apps on the Reddit platform”, as in embedded into the Reddit service itself, not mobile apps
>
> tl;dr: Paid API coming.
Pot-Kettle!
The elephant in the room but everyone is forgetting is - how much does Reddit pay its users for content? Reddit's value comes from its users, which is completely voluntarily contributed lol.
Reddit wouldn't exist without the work of volunteer moderators, as ripe for abuse as the positions are.
Even search engines provide value in that they provide alternative search functionality.
I have to say I’m not a fan of reddit, but you could also ask the question: how much do users pay to access reddit?
The web has made a lot of people feel entitled to free (high quality) services. But as developers we know building and maintaining services like reddit is not cheap (let alone free).
Why would I, as someone who’s made tens of thousands of comments, care if someone scrapes and reuses my comments. I don’t want them to pay up.
This is a really rich comment from a company that relies entirely on user submitted content and has never “paid up.”
Why would I, as someone who's made tens of thousands of comments, be happy with a corporation scraping my content to create a service that they'll turn around and charge me for? I want them to pay up, so that Reddit, this wonderful service that has given me thousands of hours of entertainment and education, can be sustainable and grow.
> This is a really rich comment from a company that relies entirely on user submitted content and has never “paid up.”
Most redditors will agree that they get much more from Reddit than they give. I for one am very happy with the arrangement I have with Reddit.
99.999% of redditors will never ever have to deal with a Reddit employee. Cronyism? What the hell has that got to do with my consumption of and participation on Reddit?
I've been a mod for a decade and never had a problem with admins.
Regardless, the attitude that they're "god" is a weird way to put it. They randomly IP banned me for calling out an admin's publicity issue during an April fools event. That's cronyism. I've used Reddit for 12 years and engaged in conversations in good faith for years. Paid for subscriptions most of that time. All relationships, business or otherwise should be mutual in some form.
https://www.redditinc.com/policies/user-agreement-april-18-2...
> When Your Content is created with or submitted to the Services, you grant us a worldwide, royalty-free, perpetual, irrevocable, non-exclusive, transferable, and sublicensable license to use, copy, modify, adapt, prepare derivative works of, distribute, store, perform, and display Your Content and any name, username, voice, or likeness provided in connection with Your Content in all media formats and channels now known or later developed anywhere in the world. This license includes the right for us to make Your Content available for syndication, broadcast, distribution, or publication by other companies, organizations, or individuals who partner with Reddit. You also agree that we may remove metadata associated with Your Content, and you irrevocably waive any claims and assertions of moral rights or attribution with respect to Your Content.
That's just about every aspect of "ownership" I can think of, minus the label "ownership". Honestly, it seems about as close as a lawyer would allow a company governed by Section 230 to have, as "ownership" would step into exposure to liability.
I think charging for the api is bad as it will make things like user apps harder. I think Reddit’s app is bad, so other apps need to use the api in order to function.
User submitted content and moderation.
It's really odd to call it "their" data, and this is not exclusive to Reddit.
They provide the platform for free. Don't like it? Self-host or go elsewhere. This is the biz model every content silo uses.
It is quite the cesspit and always has been.
Training much on it will likely worsen the confidently incorrect problems.
It's a daylight robbery. The sum of 18 years of Reddit is an enormous capital investment as well an immeasurable amount of hours spent by its users to create the content.
It's absolutely baffling how a single entity (OpenAI, Google Bard) can just take it all without permission or compensation, and then centrally and exclusively monetize these stolen goods.
The fact that we barely even blink when this happens, and that founders confidently execute on an idea like this, tells you everything there is to know about our industry. It doesn't even pretend to do good anymore. Anything goes, really.
Anyway, get ready for an "open" web that will consist of ever more private places with ever higher walls. Understandably so, any and all incentive to do something on the open web is not only pointless now, it actively helps to feed a giant private brain.
First, Stack Overflow contributions are licensed under Creative Commons. So monetizing them is explicitly allowed.
Second, information is not "stolen" nor "goods". Copyright law is completely separate from physical property laws, so even if you could make a case about fair use of training data, copyright-ability of model weights and AI generated content (which I agree are still legal gray areas) and therefore whether or not the "Share-Alike" CC clause is enforceable in this context, it would be an entirely different argument from whether the whole industry is somehow entirely morally bankrupt.
Third, given that this is unpaid work made voluntarily by users of the platforms (Reddit, SO), why is it any more acceptable for these platforms to lock it up and monetize it than for AI companies?
I think it's completely reasonable to charge for API access, particularly above a certain volume, but not because these companies have a right to protect some sort of "intellectual capital investment", but rather because the server costs of processing the requests are not negligible.
If anything, this situation really separates the wheat from the chaff in terms of what pools of open web content are truly "open". If the platforms hosting them expect to retain control of their "investment" can they really be said to be open?
I understand the irony, given that OpenAI's own name is somewhat at odds with its practices (of merely providing open access versus truly releasing everything as open source) but I think the reasonable solution to that conundrum is something like Wikimedia Foundation, Internet Archive or maybe CERN for AI, not giving up on free, open content just because it might feed a giant private brain.
Does Microsoft not cite SO posts in Bing results? Do they not make it easy to find the "correct" SO question/answer?
Is the issue that someone else is helping others, vs "you" or the "SO community"?
Incentives for free contributors (SO users) to write up good questions, good answers and debate and to come up with and vote on better solutions in the comments is to get points, recognition and yes to help others and get credit for it in their name, even though this credit is not monetary.
If Microsoft regurgitates my answers (just using me as an example, there are infinitely better contributors) without sending traffic to the SO proper website and without people voting for my answer or participating in debates and discussions on SO website proper - and in many (if not most) cases there is no single smash-hit answer and things need to be worked out and voted on - then my motivation as an SO contributor drops to a complete 0. Basically, no reason to contribute at all, since Microsoft is going to grab my answers for itself and collect the subscription (in case of ChatGPT and Copilot), and eventually the inevitable ad revenue from majority of Microsoft and ChatGPT users never leaving the Microsoft properties and never contributing to the original SO activity.
Of course, there are tons of problems inside SO proper currently as well, but none of them destroy any motivation to contribute as third-parties scraping, regurgitating the original content and keeping the traffic to themselves.
The evolution of any human legal system can be described as follows.
1. Hey guys, here is a simple set of rules we have agreed upon, to make sure there are no conflicts. Please follow them in good faith.
2. 95% of people follow both the letter and spirit of the agreed rules.
3. Some bad actors come in and only comply with the letter of rules, hacking and exploiting the system to their obscene advantage.
4. The complexity of the rules is increased to shut down the bad actors. The new rules increase costs for everyone, good and bad actors.
Repeat steps 2-4 continuously till the system is completely broken and we are all much worse off. The bad actors, "We did nothing wrong, we followed the letter of the law."
4.5. Everyday people are incited to argue about distracting, trivial issues while systemic problems snowball.
1) Copyleft licenses
2) Abolish copyright law
I am one of the few arguing for #2, but I think #1 is a good short term option.
When a model trains over Reddit, it may still provide a service that is free. But the way it's going, companies are charging money for access to those models and aren't generating traffic for the underlying training data/sites.
But make no mistake, the secret sauce in Google Search is by no means open, and possibly not even comprehensible to a single human at this point.
If you take a concept like "fair use". Let's say I embed your photo and express an opinion about it. That's what fair use was designed for. In-context relatively harmless usage of the content of others, for the sake of expression, culture and education.
That's not the same thing as "let me suck up all content ever created without permission, attribution or compensation, mangle it and sell it via the backdoor whilst making you obsolete".
You can't call that fair use, they are wildly different usages at wildly different scales with wildly different impact.
We need a new copyright category specifically for AI usage. If nothing is expressed, no training permission is given. One can opt-in and allow for training, allow for training under conditions, etc.
I work in ML so I'm aware of the consequences but society wasn't.
My step-daughter is finally crushing it as an graphics artist and she is really pissed at tools like Midjourney.
I asked her about it and she said "yes, they steal the artwork of real artists and generate fake knockoffs" ... and I don't think her opinion is invalid.
In addition, we're all kind of forced to hop on to AI whether we're a programmer or artist just to buy ourselves a little more time, delaying the inevitable. Actually, perhaps accelerating the inevitable by contributing to it.
Even in an utopian world where we would have an economic model to support this (UBI), the outcome still sucks. It wipes out human culture. There's no point in creating/producing anything as almost anything can be produced by anyone, at incredible quality, at no cost and with little skill.
Hence, your daughter being or becoming an incredible artist would have no meaning, except perhaps for herself enjoying the process of creating art.
There are lots of points and arguments to be made in this general area, but I have to ask, is this really so bad? I mean, what is the point of our lives and everything we do, other than to generally spend the rest of our time doing things we enjoy for their own sake?
If we're comparing "your daughter is an incredible artist, and here's a job for her designing product packaging for a multinational conglomerate" to "your daughter is an incredible artist, and the multinational conglomerate is using a diffusion model to design their packaging", I think it's really hard to say that the former is better than the latter. Of course, it all depends on the economic model, but the line I am quoting is within that assumption you made of the economic model being able to support this. In that case, I am for the latter wholeheartedly.
Economic incentives are great to get people "hustling", but they are rarely aligned with the human values you wish to protect, and mostly by chance if they are. Your daughter's artistry is better "spent" on art for art's (and personal enjoyment's) sake than on drawing clip art for an obscure HR form somewhere, IMO.
I know several people that without external force (work, duty) would have absolutely no idea what to do with themselves. Even their free time they organize around work-like chores or spend it on passive media.
These people seem to lack any sense of wonder, of curiosity or exploration. And it seems a permanent and fixed state. This is who they are. You can't change it.
I would not worry about this problem though because surely in the hypothetical situation of no commercial work, there's plenty of other work we can make up.
But it's only half the story. Besides the process of creating art in itself being rewarding, the other rewarding part should be how other people relate to it.
One might have trained themselves for thousands of hours and this will be reflected in the output. Most people suck at art thus the skill, dedication and creativity are recognized as such. This system has merit and scarcity.
The new system has no merit as any fool can type in a few words. Nor does it have scarcity which means an overabundance of output. Both contribute to a lost sense of meaning in creating and even consuming art.
If tomorrow we will all be as fast as the fastest runner, running will become quite pointless. There is no reward or recognition for running fast. In fact, you can't even call it fast anymore, as anybody can do it.
Does this imply that some significant portion of art "value" is derived from scarcity (e.g. there is more value to creating/producing art when a smaller portion of the population can do so)?
From a strictly financial sense that makes sense, but it does seem morally at-odds with anything that makes art easier for humans to produce.
Is it "good" or "bad" to enable a larger population to produce more art?
Is it "good" or "bad" to enable a larger population to produce higher quality art?
Culturally, both seem like they'd be good. In our current economic model, they're probably both bad.
With an economic model that supports artists financially and removes the need to transmute "art" into "money", I don't think we'd see human culture wiped out. Without a financial incentive to create art, what's the point in creating/producing anything if not to contribute to human culture?
Skill: if the merit part is entirely lost, surely we will value art far less compared to now. Anybody can make anything so what is the point?
Output: lots of art to admire is great, but unlimited art isn't. You can't attach value to unlimited.
If I thought that AI art would allow almost anything to be produced by anything at incredible quality, at no cost and with little skill, that sounds like a Sci-Fi utopia to me, an almost unimaginable world in which all limitations on self-expression are lifted. A world in which making a movie or a TV show or a video game becomes a weekend project. It sounds wonderful.
This is the most meaningful reason for creating art. In fact, I'd argue human expression is the defining element of art (AI output not being art in that sense of the word), and economic motivations just pervert it.
We’ve already been through this with cameras, which are technically just the same as using your eyes and your memory. Yet both legally and morally we all feel that operating a camera doesn’t grant you the same rights as you have by just being and looking. Strolling through the park and seeing the kids playing is very different from bringing a zoom lens and a camping chair.
That said, society could agree to a fair use that applies to ML-trained models. It could simply cover all non-commercial applications, or at the very least research.
But likely people remember references to certain art and can look them up and then 'copy paste' stylistic elements (but with a lot of effort!)
I believe the cream will still rise to the top, and the best artists will still create something totally different, and/or use AI tools to generate something better than they could create otherwise.
> I asked her about it and she said "yes, they steal the artwork of real artists and generate fake knockoffs" ... and I don't think her opinion is invalid.
Creativity doesn't exist in a vacuum. New creations are based on long-term absorptions of existing concepts & discoveries, & the decision to advance or rebel against any combination of said concepts & discoveries.
The nature of the work will change to focus more on the final product, wherein humans still hold an advantage over art generation models in terms of errors in the produced artwork. There's the possibility that such errors will be corrected with the use of an additional model down the pipeline that's solely focused on correcting said errors, but they're not foolproof either.
There will also be a larger emphasis in some niches over the documentation of the creation of said artworks, as it currently exists in some niche circles I'm in. Reductively, it's the knockoff Gucci handbag problem, wherein the remedies towards it will be the same here:
- (Tech) Serial imprinting / rollover keys / embedded signatures for verification
- (Social) Shaming & ostracization of individuals that buy knockoffs
I'm hesitant on using the legal system to solve such a problem, as the way the current copyright system is set up, it makes it near impossible for a new artist to NOT step on an existing artist's style in some form or another, even if unconsciously doing so.
https://www.stephankinsella.com/paf-podcast/kol236-intellect...
Totally agree, no question about that. But data comes from users. Shouldn't they also get paid?
I remember the good ol' days before reddit, where every community had its own forum, ran for the community, not for profit. Sure, they were running some ads to keep the lights on, but those were non targeted ads, just generic stuff based on the community.
With Dpreview dying along with so many other forums that used to serve various communities on the decline, Reddit becoming the one-stop-shop for all communities is the worst possible outcome.
It's not money that's failing, it's the rule-masters. Agents can't work well in systems with bad rules.
Money itself doesn't fail to incentivize behavior. Rather, it is what you choose to reward with money that has be carefully chosen to incentivize the behaviors you want to encourage (via monetary reward).
This is a company taking user contributions as their own, aggregating it, and using their work for free to make money off it.
As if this behavior isn't already rampant.
It'd turn into a world where people would try to make money (which still happens but normally is sniffed out), instead of a place where people like LundgrensFrontKick produce content, for free, because they love doing so, like:
* Estimating how long it took The Joker to set up the giant cash pyramid in The Dark Knight
* Comparing the box office success of movies that have a snowmobile action scene vs those that have a jet ski action scene
* Objectively trying to determine which Fast and Furious movie was the fastest and most furious
https://www.reddit.com/user/LundgrensFrontKick/?sort=top
Yes I know he has like a podcast now or something, but that only came after years of doing this for no reason other than he enjoyed doing it.
Which is today's right-wing billionaires and their pet politicians.
I'm not sure exactly how they're monetizing (maybe they sell the accounts once they have some popular posts?), but they definitely are.
Because reddit focuses mainly on what's happening recently the good content of the past that might be relevant to a user today is buried. Reposters play a valuable role in resurfacing content. I think a better paradigm, though one I can't really imagine that well, would remove the need for reposters by automatically showing the content they would repost. Maybe a recommendation algorithm?
Reddit didn't get rid of r/hailcorporate on accident. There are literal industries that exist to make fake accounts, karma farm, and sell use of those accounts to post basically sponsored messages that maybe even reddit itself doesn't know are sponsored. Think of how many people say "I search reddit for product recommendations" and know that companies have been pushing on that button for years and years. Whether reddit is honestly trying to prevent this kind of stuff doesn't actually matter, because as long as real moderation costs money and breaking that moderation makes money, the advantage is towards those who break it. FFS, reddit still has most popular subreddits modded by one account and their sockpuppets.
https://www.redditinc.com/policies/user-agreement
> When Your Content is created with or submitted to the Services, you grant us a worldwide, royalty-free, perpetual, irrevocable, non-exclusive, transferable, and sublicensable license to use, copy, modify, adapt, prepare derivative works of, distribute, store, perform, and display Your Content and any name, username, voice, or likeness provided in connection with Your Content in all media formats and channels now known or later developed anywhere in the world. This license includes the right for us to make Your Content available for syndication, broadcast, distribution, or publication by other companies, organizations, or individuals who partner with Reddit. You also agree that we may remove metadata associated with Your Content, and you irrevocably waive any claims and assertions of moral rights or attribution with respect to Your Content.
Powertripping basement dwellers who ban anyone who refuses to worship their supreme authority are one of the worse aspects of Reddit.
This applies doubly so if it's an established region based subreddit i.e. city, state, province, or country, and IMO these are the most problematic subreddits for overmoderation. Finding non-partisan regional subreddits is damn near impossible.
That's not at all what I'm saying. There are a few subreddits with fair mods who enforce the rules fairly, but the great majority doesn't - they are the rules, and if they don't like you, tough tiddies. Making it effectively a "mod and minions", not a real community with real rules.
Unfortunately, many of them start thinking, like you, that the power to moderate means they are the supreme authority and that the subreddit is about them - so they behave accordingly, feeding their ego at the expense of a community. Of course, if you believe the power itself gives them ownership over a community, that's fine. It's just that I don't.
Access to the rest of the current community.
AI companies are betraying basic business principles: they are taking value from datasets like Reddit and Medium without giving any value back. Fine if you can get away with it. But since AI, especially text based LLMs, relies on source material, it's pretty straightforward for the platforms that host that source material to deny access. Things like ChatGPT do need current source material.
I don't think it'll come to a war though and that the AI companies will instead give some value back. It could be as simple as citations that send traffic back. That's essentially the exchange of value that we all have with Google these days.
But if it's money, then I think the obligation is for platforms to pass that on the authors. It'd be hard for an individual author to negotiate this on their own with a company like OpenAI, but platforms are in a good position to negotiate on their behalf.
The AI is definitely giving value back.
I’m curious how a content platforms TOS will matchup against a Search Engine’s webcrawler TOS.
“I want people to find and access my content, but I own it.”
vs
“I will send people to your content, but any public data I can access, I can store and process how I want.”
Also, Medium has a metered paywall already. Why not just let them open up a corporate account and pay to access paywalled content the same way users do? Why are any negotiations required?
BTW I use Medium but I never use the paywall. I'm fine with my content being used to train AI for free. The payments and tax complexity involved aren't worth the tiny amount of income that any such deal might generate, nor do I want OpenAI to have a monopoly.
Of course it’s fun to watch a turf war, and we can all cheer for our favorite team and quibble about who deserves a punch in the gut.
But, we also need to keep an eye on the horizon. This will change the world, even and especially the spaces that we currently rely on. Just look at what happened to legacy media when the aggregators came: it largely turned into blogspam and clickbait. Comment sections (like this one) aren’t perfect, but they’re a damn good pressure valve for regular people to interact with the world. What will happen to those, for instance?
That being said, I don't agree about Reddit comment quality... its just generally not horrible, with the better part of it being old or in niche subs on niche topics (like fandoms or memes) that the LLM trainers are avoiding anyway.
I am concerned with the opposite, that LLMs will shout into our comment sections. For instance, building a convincing sentiment manipulating bit network will be dirt cheap and easy. Even here on HN there’s a financial incentive to flood the place with bots to promote tech products.
> I don't agree about Reddit comment quality... its just generally not horrible
That’s fine, but the important thing is that most comments are written by real people who wasted time to write it.
Not just that, but language is used as a marker. You can tell when someone talks about a subject they know a lot about. This ability to judge for yourself, based on the content alone, will be eroded. Anything can sound convincing, even to the trained ear. This makes content-oriented communities like Reddit and HN particularly vulnerable.
James Madison wrote about it, saying the future owes deference to the past by carrying on the benefits it inherits from it.
I wonder if people could just social in place more; talk, make art, rather than stare at a glass obelisk all day should social media die. You think that’s ever happened in human history? I dunno.
Third-party apps don't show ads; there's no reason ads couldn't be included in the feed and required to be shown as a condition of using the API, but I imagine it makes tracking impressions etc far more difficult. Any new features they add also need to either be incorporated into the API or remain unavailable for those users.
My only hope is that third-party apps remain niche enough that Reddit leaves them be; the first-party experiences are all awful to the point where I would probably just stop using Reddit if third-party offerings become unavailable.
https://www.reddit.com/r/apolloapp/comments/12ram0f/had_a_fe...
* Edit *
They also might be pulling a Tumblr. I really hope they don't.
> For NSFW content, they were not 100% sure of the answer, but thought that it would no longer be possible to access via the API, I asked how they balance this with plans for the API to be more equitable with the official app, and there was not really an answer but they did say they would look into it more and follow back up. I would like to follow up more about this, especially around content hosting on other websites that is posted to Reddit, as well as different types of NSFW content (a text post marked NSFW due to a gory moment in a story, for instance).
As noted in the comments, the API changes will also affect the quick .json representations of Reddit pages, which were an easy way to play with real-world data for beginners learning coding/data science.
Had no idea this existed... replacing the slash at the end of the URL and appending .json will output the post in JSON. Quite nice!
Soon Reddit's users will want a cut of it for content they create.
Then all the places these users are copying content from will want their share.
There's no solution here. Either the web stays (mostly) open and free-for-all like it is now or everyone sets up their own little walls and ends the party.
The Open Access Movement has fought valiantly to ensure that scientists do not sign their copyrights away but instead ensure their work is published on the Internet, under terms that allow anyone to access it.” - Aaron Swartz
The irony here.
A major title change came from the New York Times source that is "Reddit Wants to Get Paid for Helping to Teach Big A.I. Systems". Now that makes it much more clear what this is all about and why it is happening right now.
They are going to charge for API access, but it will remain free for more limited purposes.
* Offering an API is expensive, third party app users understandably cause a lot of server traffic
...
* To this end, Reddit is moving to a paid API model for apps. The goal is not to make this inherently a big profit center, but to cover both the costs of usage, as well as the opportunity costs of users not using the official app (lost ad viewing, etc.)
* They spoke to this being a more equitable API arrangement, where Reddit doesn't absorb the cost of third party app usage, and as such could have a more equitable footing with the first party app and not favoring one versus the other as as Reddit would no longer be losing money by having users use third party apps
* The API cost will be usage based, not a flat fee, and will not require Reddit Premium for users to use it, nor will it have ads in the feed. Goal is to be reasonable with pricing, not prohibitively expensive.
* Free usage of the API for apps like Apollo is not something they will offer, and thus me offering free usage of the app will likely be very difficult, Apollo will almost certainly have to move to an Apollo Ultra only (AKA subscription) model
...
Pulling data off Reddit now will likely give you a very large amount of polluted data from LLMs. I mean, yea it could be useful for some broad topics at this point, but still likely to contain a lot of GPTs own feedback.
It's likely that companies like OpenAI will just use their old reddit dataset, and then move to scraping things like YouTube for not just text, but audio and imagery too.
One of the best parts about social media is watching swarms of people who know nothing pivot around things you know something about.
The issue I have with Twitter's new API pricing is it's not either - it's paying a lot for a little, so feels more like an explicit move to stop companies building on Twitter. Like it's trying to kill the API altogther.
Or is it because of things you know something about and I don't that I cannot understand what you are talking about?
Your unneeded hitler comparison doesn't really inspire confidence in your knowledge either.
Do you have sources about the people comparing Musk to Hitler?
If this also extends to independent third-party clients then that's basically going to be the end of Reddit.
I think we all know that it's more column B than column A.
And while I'm not entirely comfortable with LLMs consuming all of that content without reimbursing the creators of that content. I don't see how Reddit charging for its API is different on any meaningful level.
They don't need to as they claim a royalty-free license [0] over all content posted to reddit (Section 5).
Nonetheless, markets don't operate efficiently when people horde shit.
We are the Borg. Lower your shields and surrender your ships. We will add your biological and technological distinctiveness to our own. Your culture will adapt to service us. Resistance is futile. Please upvote.I’m pro human species so I just want us to win.
The latter point isn't a bug it's a feature. Reddit is designed to function that way. The owners/execs have never expressed any interest in countering it outside of the limpest lip service imaginable.
That said, on subreddits I see people who post content without attribution all the time. I recall in /r/aww you can't directly link to an Instagram post but you can "steal" the image and post it, and it's optional as to whether or not you link to the Instagram post within the comments. Likewise, people take videos from YouTube/TikTok and re-host it on Reddit.
In smaller subreddits people will post entire pay-walled articles as if writers only get paid in likes.
You'll have an account like "Science is amazing" or something similar which seems uplifting and does show relevant/great content. Given the positive name and quality content, they get popular quickly.
But they never attribute or give back. They gain millions of followers whilst the original creators of the content get left behind. One of many things broken on the internet.
If that holds up legally, the best you can do is to try to stop your content from being scraped or not release it at all.
The right thing for them to do morally, would be to implement content visibility/privacy controls for their users similar to what Facebook offers (strange feeling to be referring to Facebook in this context).
Basically what I want is that all models trained on open source data or user created content without proper licensing are also open source and free.
(Also when the Twitter APIpocalypse happened, which the article forgot to mention.)
> Reddit is moving to a paid API model for apps. The goal is not to make this inherently a big profit center, but to cover both the costs of usage, as well as the opportunity costs of users not using the official app (lost ad viewing, etc.)
And every interaction from users with ChatGPT is valuable content provided to OpenAI.
Most people don't realize this, but every question contains information. When a user asks "Which city is better for digital nomads, Berlin or Lisbon?", they have given out a bunch of information. That there is something called "digital nomads". That there are cities called "Berlin" and "Lisbon". That those seem to be considered good for "digital nomads".
And even more so when the chat continues. If ChatGPT praises how nice a city is for studying and the users replies "I don't study. I need a cheap apartment with fast internet", the user provided information about the preferences of "digital nomads", that apartments can be cheap or expensive, that apartments have internet, that internet can be faster or slower.
No, Agents that can query current information do not fix these issues.
Folks are drastically underestimating the "grey goo" problem when it comes to training data. Now that AI generated content is so cheap to generate, the quality of training datasets is going to plummet.
Socially speaking, perhaps not so much.
> Funny timing, given the post yesterday and my praise for how communicative Reddit has been, but today there's a comparatively much more vague post about changes to the Reddit API.
> I posted in that thread and asked a few questions which as of the time of posting have not been answered.
> Shortly after the post they emailed me about a meeting, which I've replied to and will keep you all in the loop on.
> - Christian
https://old.reddit.com/r/apolloapp/comments/12qxo6l/reddit_t...
EDIT: I guess it’s safe.
> Reddit’s API will remain free to developers who want to build apps and bots that help people use Reddit, as well as to researchers who wish to study Reddit for strictly academic or noncommercial purposes.
because reddit don't actually delete anything (plainly visible in GDPR dump)
a single letter from a solicitor to his hosting would probably shut that down
Maybe they should lock their entire site behind a paywall.
There was also plenty of outrage for the new API rules, including specifically blocking weather alerts (which Elon lied about restoring): https://mashable.com/article/twitter-exemptions-nws-public-s...
Which makes the comment much much weirder since there is plenty of outage on Reddit itself about it.