[Edit] I should add that nobody has paid me for accessing my site that has a ToS/AUP requiring everyone to pay me 10% of their monthly income.
[Edit] I should add that nobody has paid me for accessing my site that has a ToS/AUP requiring everyone to pay me 10% of their monthly income.
User-agent: *
Disallow: /edit: even Twitter don't really have much faith in those rules:
# Every bot that might possibly read and respect this file
# ========================================================
User-agent: *
Disallow: /
"might possibly". lol. Really sounds like "probably won't work but can't hurt, eh?"Up to the lawyers and the courts. And I can definitely see how it could be interpreted to be a legitimate "bots are not authorized to view this site" directive since robots.txt is a standard which is respected across the industry.
I sure would hope it would not cut it in courts.
Also, you deleted your comment conveying disdain over the Twitter to X rebrand. Considering rebranding is very common in mergers and acquisitions, would you say you feel strongly about it, and if so, why do you feel that way?
I deleted it because I found this comment was crap. It did not bring anything useful, and was not funny neither. And would have risked bringing sterile discussion. It was just not at the level of what I expect from myself. Sorry if you tried to reply and failed as a consequence.
> would you say you feel strongly about it
No. There are so many things that are more meaningful and enjoyable to focus on in life that spending emotion on some rebranding is not worth it. I'm not even a user of the thing.
The rename is stupid though, even more so than my deleted comment, no doubt about it. ;-)
Those are generally accepted as legal and enforceable. Though enforcing a pants color might be hard.
Originally, the ninth circuit ruled that web scraping was allowed. This was overturned by a Supreme Court decision. It was ultimately found that HiQ was in violation of the User Agreement and they settled with LinkedIn.
> In a second ruling in April 2022 the Ninth Circuit affirmed its decision.[5][6]
Is this not saying the opposite?
HiQ no longer exists, so I think the case is definitely done. It doesn't seem to have set any firm precedent about whether web scraping violates the CFAA, but considering HiQ ultimately had to pay damages, I don't think its a resounding win for web scrapers.
[1] https://www.natlawreview.com/article/court-finds-hiq-breache...
For an interesting breakdown including from a lawyer on this very site.
At least, that's the last ruling noted in the wiki article.
The SC also basically ruled that that HiQ did not have the right to scrape LinkedIn's website.
> In its second ruling on Monday, the Ninth Circuit reaffirmed its original decision and found that scraping data that is publicly accessible on the internet is not a violation of the Computer Fraud and Abuse Act, or CFAA, which governs what constitutes computer hacking under U.S. law.
Here, bartenders can say you've had too many, they are supposed to stop selling you alcohol if they think it could be dangerous for you or other people. This is one of the cases where not only the business can refuse to serve you, it also has to.
Now, I'm not a lawyer but I wouldn't expect robots.txt to be legally enforceable.
Of course, in any case, the website can totally refuse to fulfill a request, rate limit, etc. I'd say this is not comparable to the bartender or a business refusing to serve you though. Responding to an HTTP request is not offering a commercial service. And if it is, whatever legal document that promises you the service can totally specify that they won't promise to answer your robot.
To this day, bars have a list of "refuse entry". Sporting teams have life time bans to hooligans. Sure, the ban hammer should be slow to swing, but it shouldn't a never allowed to swing situation.
Were it a standard, it wouldn't be so easily neglected without consequence. In practice it's an arbitrary playground rule, like declaring the floor is actually lava. Unless you're directly involved with coming up with these "standards," you can go your entire career in total ignorance of all of them.
It's there for entities willing to (try to) respect website owners wishes and preferences.
It's useful when you need accessing stuff in an automated way, while staying on good terms.
By downloading robots.txt, the entity is asking the website owners their preferences regarding crawling. It's polite to do so. robots.txt can also offer useful guidance ("yeah, don't go there... [it's useless, too costly, etc], but you can go there...").
It's not a guarantee. And entities can totally decide that the request is unreasonable. Their call. The risk is being shamed or blocked.
Personally, I find it unreasonable / harmful to allow Google and deny everything else. If I built a crawler for a search engine, I would probably not follow this. And Twitter might find out and block me, of course.
robots.txt is not a security measure / measure against abuse. This is what IP ban and other kinds of blocking are for. But it's still useful. Not flawless, but useful.
(Twitter specifically allows Google earlier in its robots.txt file.)
A. "I'm going to scrape the web to do X, Y, Z"
and
B. "I'm going to scrape Twitter data to do X, Y, Z"
A is far more likely to care about robots.txt because there's enough left of the web to complete their mission. B will not care if Twitter has a robots.txt file.
I wonder about Bing though, since they do seem to have Twitter results.
Sounds like discrimination.
Why Google can scrape data, and I as individual can't?
Sorry I'm from Europe, so I'm not that much into "what corporate wants, it gets" kind of thing.
Yes. Views of logged-out profiles display statuses sorted by most to least likes, rather than chronologically. It's incredibly dumb, and makes many accounts (like government announcement accounts - @NWS et al) useless while logged out.
What a disaster.
Some people like(d) to compare Musk to Jobs; but when it comes to Product instincts - they couldn't be further apart. Musk is the anti-Steve-Jobs, his ability to make the worst product decisions is preternatural.
I deleted my Twitter account before all that, all too often I cannot even see Twitter deeplinks anymore.
However, if you actively try to circumvent this refusal of the service provider, this could very well be a crime, depending on what you are doing.