X updates its Terms to prohibit crawling/scraping of its data
stackdiary.com
stackdiary.com
I know it'll never happen because we live in the reality where everything must be awful all the time, but I'll continue dreaming of a day where large internet platforms are regarded and regulated as public utilities.
Tweeting is the ultimate example of an activity that's designed to be engaging but is otherwise awful. I think it may be the very nature of popularity-driven ultra-short-form content that it encourages online mobs, cancel culture, dunks and hot takes.
If there's even a chance this is true then it would be criminal to calcify this medium in its currently known form as a permanent pillar of our public square - which regulating it as a public utility would do. We could be inflicting suffering on future generations that might have escaped it.
No. It's better to hope that we discover better ways of communicating, and tweeting becomes irrelevant and dies.
What "cancel culture"? Who got "cancelled" for real? Literally everyone of those whining about "cancel culture" who hasn't actually broken laws or is being suspected having done so (Tate) is still running around and making tons of money whining around to everyone.
Meanwhile innocent people on the left get their lives ruined by cancel culture twitter zombies that think harassing and stalking people makes them righteous. And just the usual, people who want to be assholes and just pick the 'socially acceptable' way to hurt other people. "The Left" is extremely cannibalistic among those who subscribe to a central "Left" dogma.
There's unfortunately massive amount of content produced and reproduced by people who imagined they were doing so on a public forum. But it's increasingly going behind a wall of 403s and obfuscation.
For a few days in July non-Twitter members couldn't see anything there. Now they can only see the primary post and not the threads. That's Musk's right to do that, but the right thing to do is for people to clue in that this isn't a public resource, and if they want their content to be broadly available they'll need to find other places to put it. And not just for posterity, but for the present as well.
One thing I don't understand is how the advertisers who nominally fund Twitter (or did) are fine with this. Seems like deliberately kneecapping the reach and impressions.
Of course, that's how settlements work. Elon literally has more wealth to bully people than anyone else on the planet, and is willing to do things out of spite.
[Edit] I should add that nobody has paid me for accessing my site that has a ToS/AUP requiring everyone to pay me 10% of their monthly income.
Yes. Views of logged-out profiles display statuses sorted by most to least likes, rather than chronologically. It's incredibly dumb, and makes many accounts (like government announcement accounts - @NWS et al) useless while logged out.
What a disaster.
Some people like(d) to compare Musk to Jobs; but when it comes to Product instincts - they couldn't be further apart. Musk is the anti-Steve-Jobs, his ability to make the worst product decisions is preternatural.
I deleted my Twitter account before all that, all too often I cannot even see Twitter deeplinks anymore.
User-agent: *
Disallow: /edit: even Twitter don't really have much faith in those rules:
# Every bot that might possibly read and respect this file
# ========================================================
User-agent: *
Disallow: /
"might possibly". lol. Really sounds like "probably won't work but can't hurt, eh?"Up to the lawyers and the courts. And I can definitely see how it could be interpreted to be a legitimate "bots are not authorized to view this site" directive since robots.txt is a standard which is respected across the industry.
I sure would hope it would not cut it in courts.
Also, you deleted your comment conveying disdain over the Twitter to X rebrand. Considering rebranding is very common in mergers and acquisitions, would you say you feel strongly about it, and if so, why do you feel that way?
I deleted it because I found this comment was crap. It did not bring anything useful, and was not funny neither. And would have risked bringing sterile discussion. It was just not at the level of what I expect from myself. Sorry if you tried to reply and failed as a consequence.
> would you say you feel strongly about it
No. There are so many things that are more meaningful and enjoyable to focus on in life that spending emotion on some rebranding is not worth it. I'm not even a user of the thing.
The rename is stupid though, even more so than my deleted comment, no doubt about it. ;-)
Those are generally accepted as legal and enforceable. Though enforcing a pants color might be hard.
Originally, the ninth circuit ruled that web scraping was allowed. This was overturned by a Supreme Court decision. It was ultimately found that HiQ was in violation of the User Agreement and they settled with LinkedIn.
> In a second ruling in April 2022 the Ninth Circuit affirmed its decision.[5][6]
Is this not saying the opposite?
HiQ no longer exists, so I think the case is definitely done. It doesn't seem to have set any firm precedent about whether web scraping violates the CFAA, but considering HiQ ultimately had to pay damages, I don't think its a resounding win for web scrapers.
[1] https://www.natlawreview.com/article/court-finds-hiq-breache...
For an interesting breakdown including from a lawyer on this very site.
At least, that's the last ruling noted in the wiki article.
The SC also basically ruled that that HiQ did not have the right to scrape LinkedIn's website.
> In its second ruling on Monday, the Ninth Circuit reaffirmed its original decision and found that scraping data that is publicly accessible on the internet is not a violation of the Computer Fraud and Abuse Act, or CFAA, which governs what constitutes computer hacking under U.S. law.
Here, bartenders can say you've had too many, they are supposed to stop selling you alcohol if they think it could be dangerous for you or other people. This is one of the cases where not only the business can refuse to serve you, it also has to.
Now, I'm not a lawyer but I wouldn't expect robots.txt to be legally enforceable.
Of course, in any case, the website can totally refuse to fulfill a request, rate limit, etc. I'd say this is not comparable to the bartender or a business refusing to serve you though. Responding to an HTTP request is not offering a commercial service. And if it is, whatever legal document that promises you the service can totally specify that they won't promise to answer your robot.
To this day, bars have a list of "refuse entry". Sporting teams have life time bans to hooligans. Sure, the ban hammer should be slow to swing, but it shouldn't a never allowed to swing situation.
Were it a standard, it wouldn't be so easily neglected without consequence. In practice it's an arbitrary playground rule, like declaring the floor is actually lava. Unless you're directly involved with coming up with these "standards," you can go your entire career in total ignorance of all of them.
It's there for entities willing to (try to) respect website owners wishes and preferences.
It's useful when you need accessing stuff in an automated way, while staying on good terms.
By downloading robots.txt, the entity is asking the website owners their preferences regarding crawling. It's polite to do so. robots.txt can also offer useful guidance ("yeah, don't go there... [it's useless, too costly, etc], but you can go there...").
It's not a guarantee. And entities can totally decide that the request is unreasonable. Their call. The risk is being shamed or blocked.
Personally, I find it unreasonable / harmful to allow Google and deny everything else. If I built a crawler for a search engine, I would probably not follow this. And Twitter might find out and block me, of course.
robots.txt is not a security measure / measure against abuse. This is what IP ban and other kinds of blocking are for. But it's still useful. Not flawless, but useful.
(Twitter specifically allows Google earlier in its robots.txt file.)
A. "I'm going to scrape the web to do X, Y, Z"
and
B. "I'm going to scrape Twitter data to do X, Y, Z"
A is far more likely to care about robots.txt because there's enough left of the web to complete their mission. B will not care if Twitter has a robots.txt file.
I wonder about Bing though, since they do seem to have Twitter results.
Sounds like discrimination.
Why Google can scrape data, and I as individual can't?
Sorry I'm from Europe, so I'm not that much into "what corporate wants, it gets" kind of thing.
However, if you actively try to circumvent this refusal of the service provider, this could very well be a crime, depending on what you are doing.
"Googlebot" might as well be the "Mozilla/5.0" of search index bot user agents.
What you mean by that? Currently it's not enforceable by any technical means, as I can scrape Twitter without any limitations.
if(requests > bigNumber) banHammer();
Or whatever alternative bot detection they want to do. There are plenty of solutions for these types of problems.If requests from who? or what? Account? IP? Because no one is doing it from one account or one IP.
> There are plenty of solutions for these types of problems.
Actually not a lot, it's a really hard problem to solve, unless you're referring to unsophisticated scrapers that have no idea what they're doing and are following some online tutorials on how to scrape Twitter. Out of desperation, they even implemented blanket bans on entire legitimate mobile carriers in some EU countries, but it didn't have a lot of impact on the business of Twitter scraping.
Yes, it is a hard problem. Legal justice is also a hard problem. My above point is that twitter can unilaterally take steps to enforce their policy, i.e. they do not need to lean on legal mechanisms.
Non-business-entities of course may escape judgement if they cannot be found or are in a place that doesn't enforce these kinds of laws. However, I bet Twitter's bigger fear is the US and EU commercial entities and generative AI - limiting access and then going after them in court will stop a lot of commercial enterprises.
https://www.natlawreview.com/article/hiq-and-linkedin-reach-...
https://www.natlawreview.com/article/court-finds-hiq-breache...
Twitter is locked behind a login wall now so the rules in theory might be different.
Of course, bots can still get in through the human door.
1-2 years ago they used to automatically suspend new accounts without a phone number within 1h after registration; after that you had to go through a manual account recovery process. Did something change here?