Videolan.org robots.txt
videolan.org
videolan.org
Correct me if I'm wrong: After the recent web scraping ruling[1] it seems that it's perfectly legal to ignore the robots.txt.
https://www.turnitin.com/robot/crawlerinfo.html
> Q: How can I completely exclude TurnitinBot from my site?
> To exclude TurnitinBot from all or portions of your site all you have to to do is create a file called robots.txt and put it in the top most directory of your web site.
I don't know how well this would work with a CDN, but presumably if you pay for the right tier of Cloudflare (or whatever) you can perform similar operations to prevent content being hoovered from their by clients you'd prefer not to serve.
But my favorite robots.txt is,
User-agent: Zombies
Disallow: /brains User-agent: Zombies
Disallow: /braaains
?But yeah, maybe not a good fit for this list of educational and copyright parasites.
1. So that case was about the CFAA (Computer Fraud and Abuse Act). So at most it would say that ignoring the robots.txt does not violate the CFAA -- a law that makes some things felonies as "hacking", basically. I agree that ignoring the robots.txt (say if you are Archive Team? [1]) should not be considered a criminal "hacking" felony.
But there can still be other reasons ignoring the robots.txt is against a law -- or cause for a civil tort action. (Most copyright violation is a civil tort action for instance, the CFAA is, again, a law that establishes some felonies with many years of jail time, intended to punish "hackers"). The decision in that case said nothing about anything except the CFAA.
For instance, taking copyrighted content from the public web and re-selling is probably still going to put you in various kinds of legal trouble -- just not a CFAA violation. It's possible ignoring a robots.txt could put you in other kinds of criminal or civil trouble, depending on the particular circumstances -- just not a CFAA violation. It would be interesting to research what other possible liability there might be. If for instance you caused harm to the site by ignoring the robots.txt (say, an accidental or intentional DOS), I bet there'd at least be cause for civil tort.
2. Even so, even under that case, if that specific case didn't involve a robots.txt (did it?), it's always possible the presence of a robots.txt would result in a differnet outcome. My sense is probably not though, that Supreme Court decision referenced by the ninth circuit on remand -- probably does mean ignoring a robots.txt is not a violation of the CFAA. (And again, I say, PHEW, that would have been terrible if it were -- if say someone trying to archive MySpace before it went away could be put in prison for a couple decades for disrespecting the robots.txt).
In this case Linkedin sent HiQ a cease and desist letter before they sued and claimed that letter revoked access for the purpose of the CFAA, so not quite the same as a robots.txt but legally it's probably close enough. If anything it's stronger because HiQ can't claim they didn't see it.
The 2020 Supreme Court decision in Van Buren v. United States that the Ninth Circuit said they were relying on has been decided though, and does seem pretty pertinent and a (welcome, thank god) nail in the coffin of CFAA overreach. https://techcrunch.com/2021/06/03/supreme-court-hacking-cfaa...
(Although now that I actually look at THAT case... damn! I think that's a WAY better CFAA case than a lot of them! I guess it took a police officer being the defendent to get the currently pro-police Supreme Court to actually say CFAA prosecution was going too far, okeydoke).
But yeah, nothing is for sure... almost ever with the law. But it does seem to be moving in a direction...
In practice, this could easily fall under fair-use, though, since the paper isn't distributed for the purpose of consumption.
Authors generally have right to their works by default, even students. Those who use copyrighted works for academic purposes do get some exceptions to copyright law.... but a commercial service is not this.
Furthermore, student assignments are additionally protected by FERPA.
I'm not sure the answer, but I do think it's a great question.
https://help.turnitin.com/Privacy_and_Security/Privacy_and_S...
So your school or university are really deciding whether your paper is stored for future plagiarism checks.
U.S. copyright law provides copyright owners with the following exclusive rights:
- Reproduce the work in copies or phonorecords.
- Prepare derivative works based upon the work.
- Distribute copies or phonorecords of the work to the public by sale or other transfer of ownership or by rental, lease, or lending.
- Perform the work publicly if it is a literary, musical, dramatic, or choreographic work; a pantomime; or a motion picture or other audiovisual work.
- Display the work publicly if it is a literary, musical, dramatic, or choreographic work; a pantomime; or a pictorial, graphic, or sculptural work.
-Perform the work publicly by means of a digital audio transmission if the work is a sound recording.
Copyright also provides the owner of copyright the right to authorize others to exercise these exclusive rights, subject to certain statutory limitations.
By that logic, isn't it the same for Google? They keep a copy of your content in their cache/index (you can even get the indexed page directly from their cache).
You can opt out of that, also it's published content and caching has a carved out exception. When you use this service at a school the teacher submits your paper to the service and they store it to validate against other students papers as well.
They don't only check against "published" sources.
So, sure - download to your heart's content. Which means TurnItIn has little incentive to keep works out of their database - you can't make a claim against them if they don't publicize that your work is in it.
Combined, TurnItIn is offering a service that individual copyright holders are likely uninterested in providing themselves, the use does not reduce the marketability of the copyrighted work, and there is probably not a market for this service that individual copyright holders could monetize. This is a pretty good case for fair use.
* plagiarism is not generally against the law, although it is a violation of school policies and can get you punished by your school. It's claiming someone else's work as your own. It may or may not involve a copyright violation, it can be plagiarism without involving a copyright violation -- it could be the author gave you permission, or the item isn't in copyright, or it would count as "fair use" -- it's still plagiarism if your teacher/school/honor code panel says it is.
* copyright is about the law, violating a copyright is against the law and can get you civil or criminal penalties. It involves copying work someone else legally owns without permission. It may or may not involve claiming the work as your own, for the most part whether you attribute something properly or claim it for your own is not relevant to whether it is a copyright violation. (I suppose in some edge cases it could be relevant to whether you have a "fair use" defense, but mostly it's not significant in whether something is a copyright violation).
If you provide copies of your work for free on the internet, this is why they get to keep one, just as everyone else? They are probably not allowed to distribute it, though?
A more interesting question is, if these companies do well and stay in business a long time, won't it become increasingly difficult to write an original paper that isn't flagged for plagiarism? There's only so many ways to describe the effects of the Lend-Lease Act on postwar Europe.
What I do have a problem with is the fact that, after uploading an assignment, I am required to click a checkbox that says "I agree to Turnitin's end-user license agreement." I should not have to agree to a license for a piece of software that I'm not even using; it's my professor who's using Turnitin's services. And if it's 11:55 PM and I'm trying to submit my assignment, it feels really scummy to suddenly force me to sign a legal contract that I don't even have time to read.
Then my university moved to Canvas and TurnItIn. At first there was no license agreement check box, and all the courses were force-enabled to allow TurnItIn to store student submissions forever.
I raised a lot of bell over that and the next term there was that same checkbox that I assume you also see.
It always felt very coercive. I hated checking that box. I fought tooth and nail. I had conference calls with the Academic Technologies leadership. They absolutely didn’t understand the objection. They compared it to Office 365 and didn’t understand the point that neither the university nor Microsoft was requiring that I give them a perpetual, virtually limitless license to my content in order to use the service.
I pointed to the university policies which explicitly and very clearly categorized non-compensated student output as the property of the student, who was to regain all rights. I pointed out the conflict of interest that iParadigms brings to the table.
All I ever got in response was the talking points I found on the TurnItIn marketing material. I’d have been OK if they disagreed after an actual discussion, but they weren’t interested.
IME, using a sitemap is much more efficient. For example, HTTP/1.1 pipelining can be used to reduce the number of TCP connections needed.
Is resource exhaustion what draws a public website^1 operator's attention to "bots". If it is not resource exhaustion then what is it.
1. For this question, assume "public website" means a website serving public information where there are no legitimate intellectual property rights in the information that can be asserted by the site operator.
Unexpected place to see latin1 -> utf8 mojibake
I was quite surprised to see all the weird bots that were crawling it.
># --> fuck off.
comment that is added after 3 specific robots.
Toss a coin to your admin, oh valley of plenty~~~
(I doubt they're pro-plagiarism - not even copyright abolitionists go that far.)
And yet they don't disallow Googlebot! For obvious reasons.
¹ https://www.gnu.org/software/rcs/manual/html_node/Concepts.h...
² https://lists.gnu.org/archive/html/info-gnu/2022-02/msg00001...
# --> fuck off.
User-Agent: kome
Disallow: /