The Wiki is in French (2011)
hccp.org
hccp.org
"The Spider of Doom" (2006)
A crawler deleted all of the content in a CMS, because the delete button used a GET request, without authentication, and with the URLs embedded in the HTML.
An <a> tag can only ever send a GET request. It can't be made to submit a form (without JS) afaik. The only simple way to have something the user can click to send a POST request is to use a <button> or <input type="button"> in a form, and styling that to look like a link is absolute hell.
I'm sure that if we had proper HTML tools from the start to say, "I want this thing to look like a link, but I want it to send a POST request", we'd at least have less such issues.
It’s a mess, but as far as I’m considered “not messy” HTML+CSS+JS was never an option to begin with, at least for most people.
The thought process should have been as straightforward as, "I want this link to modify state, so I'll use <a method='post'>".
I agree that using JS is probably okay for most tasks these days, but 1) a lot of web platforms were written at a time before JS was widespread, and 2) just throwing together some HTML is a lot easier than adding javascript to every link. (If you don't think that speed/ease of programming matters; you'd be surprised to learn how many internal tools various places use were just thrown together and isn't a polished product.)
This might be MS' fault, but it's also Facebook's, Slack's, pre-FB-WhatsApp's, to name just a small few of many.
<aside> On the other hand, if you're putting up a GET endpoint that's not only not idempotent, but unauthenticated and wipes your whole DB. Even on a test site. Blaming MS for your troubles is a bit much.
But wait! Surely there are other possible approaches. Can't users just report spam? Surely there are clear usage patterns that can be detected too!
These are valuable and useful approaches, but have historically proven unequal to the task. It can be tricky to spot a spam campaign spread over N accounts - only the most basic forms of spam are really obvious. URLs are easily disguised and changed. Loading the URLs and checking the contents that come back is one of the better approaches. Users also don't tolerate a bunch of spam. They will migrate away.
So, yes, anyone who makes a GET endpoint into an unauthenticated change-things URL deserves all the misery they get.
It's really not the messenger services' problem that someone's web app claims to speak HTTP but doesn't.
That is entirely his fault.
Quoting the article.
Now wondering what the wiki software was that had implemented it as a GET (of course, opening up to CSRF too if there's no token param) and implemented poor search functionality.
> one of the admins was experimenting with Nutch as a supplement to the wiki's impoverished search capabilities and had authenticated the crawler using their admin credentials
> After determining the source IP of the crawler (one of the admins was experimenting with Nutch as a supplement to the wiki's impoverished search capabilities and had authenticated the crawler using their admin credentials)
The real problem is that a GET request is meant to be side-effect free. A crawler only issuing GET requests should not be able to modify e.g. global settings. Even when using an admin token.
Meant, yes, by somebody, but a web service creator can decide otherwise, for some reasons, like simplicity.
If I remember correctly, there was a good story about Viaweb, how they figured that sending requests to follow links can be used as commands - and I wasn't sure they didn't use GET for that... but maybe I'm wrong.
As the conclusion of the article shows, if you do that you're breaking the contracts of the web platform. It's a sure way to build a web service that won't play along well with other web services, like a search crawler.
Simply introducing unexpected side effects, you mean? Using methods in ways they're not meant to be used doesn't create simplicity, it creates complexity. Suddenly there are exceptions to the standards that you need to remember and take into account. This story demonstrates that very well.
Laziness is not the same thing as simplicity.
This is like "undefined behavior" in C. Sure, you can make a GET have side effects in your application, but everything else will still assume that HEAD and GET is free of side effects, and might repeat requests, omit requests (using cached data), or even do requests speculatively in advance.
> https://twitter.com/rombulow/status/990684453734203392?lang=...
TL;DR: Safari figures out the user is visiting this page very frequently so whenever Safari opens, it tries to send a request to fetch the garage door "page" link before he visits anywhere to cache the response. Which immediately opens the garage door.