When you have financial incentive to build your business on someone’s data and you scrap literally millions if not billions of pages - it’s unethical.
This data is often of great public value. I track conversations around a social issue as part of my work for a non-profit.
I'd counter it's unethical to prevent people from accessing this data.
> great public value
Having been to twitter mostly through the most recent prominent war, man the signal to noise ratio is really low even when being careful about who to follow and who to block. There is so much disinformation, bad takes, uninformed opinions presented as facts, pure evil, etc.
So I guess it could be used for training very specific things or cataloging the underbelly of humanity but for general human knowledge it’s a frigging cesspool.
I do not use the Twitters myself, and actively discourage others from doing so. Sends people bonkers.
We were actually gearing up to switch to paid accounts as we found use cases that could subsidize these efforts... And then the starting price for reasonably small volumes shot up to like $500k/yr.
We have robots.txt. If Google doesn't respect that, it's unethical. Don't you think so?
Not that the people you want to respect that would
Especially since they're not moderating things or anything.
Agreed. However, it's probably covered by their terms of service.
Same thing with the recent reddit kerfuffle. I'd have much preferred a Usenet 2.0 instead of centralizing global communications in the hands of a handful of private companies with associated user-hostile incentive structures.
If it doesn't respect robots.txt, it is unethical.
b) Given Twitter is public, user generated content which they don't own but simply have a license I wouldn't call it unethical in the slightest.
I do a lot of data scraping, so I’m sympathetic to the people who want to do it, but violating the robots.txt (or other published policies) is absolutely unethical, regardless of the license of the content the service is hosting. Another way of describing an unauthorised usecase taking a service offline is a denial of service attack, which (again, if Musk’s description of the problem is accurate) seems to be the issue Twitter was facing, with a choice between restricting services or scaling forever to meet the scrapers requirements.
Personally I would have probably tried to start with a captcha, but all this dogpiling just looks like low effort Musk hate. The prevailing sentiment on HN has become so passionately anti-Musk that it’s hard to view any criticism of him or Twitter here with any credibility.
The only reason these websites and platforms aggregate any content at all is because they're effectively giant public squares.
Musk is trying to have his cake and eat it...
(Clearly it's not a public square, but his position is incoherent).
It is NOT legal to install cameras that record everyone's conversations, much less sell the laundered results.
Pre-2023 people went on Twitter with the expectation that their output would be read by humans.
A traditional search engine is different: It redirects to the original. A bastardized search engine that shows snippets is more questionable, but still miles away from the AI steal.
Expectations =/= reality. And the reality is that bits have been reading comments for over a decade.