Request bodies in GET requests
evertpot.com
evertpot.com
Intermediate proxies and browsers apply different caching rules to GET vs POST responses. This can have huge performance and security implications.
Along the lines of security, browsers impose stricter security mechanisms on POST requests. Cross site POST calls require a CORS pre-flight options request, and the response of this request is scrutinized. Choosing GET vs POST can be the differences between having CSRF protections and having none.
Sometimes I feel that some of the opinions here are not very useful without the context of the environments in which they are deployed. A few weeks ago I saw someone wrote that logging is an anti-pattern because correct code doesn’t need it. That may work for a small project, just like using fat GET requests may be fine for most projects. But when you are building complex software at scale, the importance of following the normal practices becomes very evident
Not in all cases. Certain POST requests count as "simple" requests and no preflight check is performed. For example, a POST with application/x-www-form-urlencoded data and no extra headers doesn't need a preflight, because you could make an equivalent request with an HTML form.
https://developer.mozilla.org/en-US/docs/Web/HTTP/CORS#simpl...
This, or having a body on a GET request, is the best option due to caching implications, of which said implications are often opinionated (not according to a spec) along the entire path of request/response.
Both PATCH and QUERY are newer and they are really specializations of what people used to use POST for.
If these are POST, you find yourself trying to conditionally cache POST requests, which opens up a whole new class of caching bugs.
1. pre-existing behavior. maybe queries have only started getting big enough to where this mattered, but all of our existing clients communicate the original way and won't migrate
2. more complex behavior on clients and services as you've now introduced a new stateful resource between client and service. should clients store query ids, or does the service handle idempotency? what db does the service use to store queries? how do we limit stored queries? do clients have to manage number of queries stored? does the service?
This is a lot of application logic to add just to avoid adding a QUERY method. It's just too much.
pretty much the same pro's and con's too
safer, more controlled, better optimized but it's more limiting, another dev dependency and a responsibility bottleneck
I did write some tracking code that ended up having really long strings and a ton of parameters and wrote it to break into multiple requests at like 1800.
I am curious about the specifics here. Do you do it at the application level or proxy level?
Does doing it at the app level require that all the requests can be routed to any DB instance regardless of whether they are master or slave. A write to a mster is passed through. A read to a master is routed to slaves LB. A read on slave is passed through. A write on replica is routed to the master.
Does this look like a sound design?
* We have an HTTP header you can set in a reply that says "replay this request elsewhere".
* We boot up clusters of Postgres servers in read-replica configurations --- a single writer, lots of readers.
* If you try to write to a read replica, you get an error from Postgres, and your framework passes that error up to you.
* We suggest our users catch the exception and set the "replay elsewhere" header to redirect the request to the write master.
And that's pretty much it. Most apps are read-heavy. Reads get serviced from replicas (and, usually, close to their users --- that's the point of the service we built) and writes go to the single write master. You don't write any serious code to make that work.
But if apps could reliably say "a POST isn't just a complicated GET, but is almost certainly changing mutable serverside state", this would be an even easier problem; you'd just have the CDN route POSTs to the write master region, and GETs to the nearest region, and you'd be done.
You can architect your apps like this today, I guess, using QUERY or abusing PUT or something. My point is just that there's significant value in being able to shift all the read-only operations out of POST.
POST, on the other hand, is expected to mutate data. It needs to go to the master node, unless you have bidirectional replication, which is way more complex.
This is why a load balancer / proxy can spread GET requests and cache their results, but can't do so with POSTs.
QUERY is like POST in that it has a body, but is like GET in that it's idempotent, and should return the same result given the same URL and body. Load-balancing and caching work again.
No, gets are safe (do not induce any client-responsibility state changes) and idempotent (do not induce any additional client-responsibility change of state for a second identical request with no intervening action after the first), not pure (do not depend on any state outside of the request).
So an HTTP framework would not retry, moving retry logic into the application.
If you can only use POST, then you also can't cache the result because POST is not safe: you've no idea if the same POST is intended to produce the same response.
Can cache in your JS code. But for browsers:
Do browsers respect cache headers for GET requests to say an API endpoint? maybe with an etag?
… Now, back to reality. :-(
Glad that a QUERY method will be provided instead, and I learnt my lesson about giving against-the-grain responses to interviewers.
Not passing judgement - you may have very well qualified your statement. But having dealt with the above issue some years back (as in being the person on the debugging run) I do sympathize with the unhappiness. :)
I've had a similar interview where I was interviewing and this came up, the interviewee didn't pass but it wasn't because they said GETs could have bodies. That was just one of the reasons, and they would've passed if that was the only "wrong" answer.
If they had explained that you could indeed have a body in a GET request, even if it went against the spec and you'd probably have to modify you existing body parser to comply, I would have accepted the answer as "correct". What matters is if you can explain your answer, not if you can answer yes or no questions.
In the end they didn't explain their reasoning and I just said that GETs can indeed have a body, but it goes against the spec, so it's gonna be hard unless "you own the entire stack", like a parallel comment said here. This way they have a base to explore more info later and improve for their next interview.
I've had plenty of intelligent answers that seemed weird to me before I followed up. Or "wrong" answers because I didn't explain the question well or there were different assumptions. Red flags shouldn't be red flags until you probe and give the candidate a chance to explain (or dig a deeper hole).
Ideally an api gateway should have managed the headers but ours didn't. Not at that time at least.
One thing I can think of: there some optimization that would cause performance degradation for 99% of GET requests.
Careful with the attack surfaces that entails, though: https://blog.cloudflare.com/understanding-our-cache-and-the-...
These discussions sometimes go off the rail because it's a bunch of STEM folks critiquing the humanities, so I'll share a STEM example that gets the point across. A too-clever friend of mine got marks off in our Analysis course because he presented a constructive proof of a theorem on a test, and that was not the proof technique that the test was designed to interrogate.
Superficially his answer was "correct". But then, as his professor pointed out, on an even deeper level it really was incorrect, because our write-ups of proofs are really high-level descriptions of formal derivations, and the student was describing a proof a different axiomatic system than the one assumed by the question, so, to be totally correct, he would have had to embed the proof that his proof in constructive analysis mapped over into normal analysis. Which would have been quite a bit of work that the student, at the time, was unable to do (and probably no one could do in an exam setting).
But, again, the real point is that this part of the debate is pedantic because what my clever "friend" really didn't understand was WHY the question was being asked. The answer is supposed to demonstrate a certain piece of knowledge in a certain context; it's an exam answer in a closed curriculum, not a Treatise.
Besides, you are ALWAYS writing into a audience and context; if you want to write to a "pure" audience, you may do so, but consider saving it for Sunday morning prayers.
Which, since that friend isn't paid to write analysis proofs these days, were perhaps important lessons ;)
Generally speaking there's nothing wrong with against-the-grain responses if they're well argued for, especially if you can compare them to the other, more common, options and make your solution sound the best. "Elastic search does this, and it might be easier to cache" don't sound like the best arguments to me.
Upcoming new HTTP QUERY method - https://news.ycombinator.com/item?id=30153995 - Jan 2022 (52 comments)
Recent:
IETF: The HTTP Query Method ( Draft) - https://news.ycombinator.com/item?id=29794838 - Jan 2022 (125 comments)
[1] https://github.com/square/okhttp
But what about DELETE with body? It's similarly unspecified, but it seems to me to be of a different breed. The point made about GET is that, under the assumption of a theoretically idempotent and immutable operation, various caching systems can kick in and serve a request from any potential world-wide server. But a DELETE is by nature a mutating event, and will always need to be directed to some kind of principal server, meaning it should probably escape the caching the same way POST requests do.
Is it just as much a gamble to do DELETE with body as it is GET with body?
That's a problem.
> But what about DELETE with body? It's similarly unspecified,
And faces the same risk that intermediate systems might reject it or strip the body, called out in RFC 7231.
> But a DELETE is by nature a mutating event,
DELETE is, more to the point, and despite being idempotent, expressly not cacheable, so there is no risk of getting served the wrong cached response based on the body being stripped. But, aside from the request not fitting the defined semantics of what DELETE means, you still have the risk of rejection or the request having the body stripped.
Now there’s this thing called REST that we all commonly understand and debate over. And I have run into so many issues that I basically abandoned calling it REST and now just call it an API.
And somehow it still all works just as well :)
You can delete the rest of the spec and clients can connect without head-aches but only if you can put up with the fanatics in your team trying to trace every bug to not following rest, in the rare case giving you no options other than trying patch after patch to make everything rest-compliant.
As long as HTTP is concerned (which must not be specifically REST), HEAD needs to be supported next to GET, too.
*except when the query string is too big, then use POST
I have personally gone off the deep end and started writing 200, 400, and 500 for all my status codes instead of the specific 2xx, 4xx, 5xx ones.
If details of an error response need to be mentioned, it will be in the error response body.
My motivation came from GraphQL as it responds everything with 200s, and so far it has been humming along fine.
True REST are more guidelines if anything.
This is slightly annoying though, as other consumers (like `fetch` in JavaScript world for example) behaves differently if there is a error. If every status is 200, suddenly you need to manually check the response inside the returned promise, while if you answer with a correct status code when the server encountered an error (5xx), then it'll jump to the `.catch` part of the promise chain instead, which you're probably already handling anyways.
Similarly with curl if I'm not mistaken (don't have it in front of me). A 2xx would make the process return 0, while a 5xx would make it return non-0, so you can handle errors without having to read the actual response.
Similar with curl, 502 returns non-zero, 302 not. 4xx IIRC as well zero unless --fail is in use. The default makes it sensible/useful for smoke tests.
Perhaps GraphQL was aware of such client side "bugs" it choose to go "OK" first of all.
I can't believe the people here saying "fuck it, I always return 200"
For example, if you're rate limited (429), a robust API will still give you back data on how many requests you've made and how many requests you're allowed to make, so you're still going to have to check the payload regardless.
The combination of specific error status codes along with error messages has been redundant in my experience.
Furthermore, in Axios, checking for `error.response.data` isn't terribly far from `error.response.status`, so I'm unsure of this "worst client experience" you're talking about.
It's pretty intuitive for me.
I think its objectively worse that I, as the engineer, need to handle all of them differently when they all could have returned to me 429 instead and I could write a general wrapper for that.
how is
if error.response.status === 429 // then do something
different than
if error.response.data.err === 'RATE_LIMITED' // then do something
Also, do you mean to tell me that you've never had to elaborate on your error messages and just sent back status codes with no body?
I'm sorry but that does not sound like a robust API to me.
90% of my error responses have specified data, and I have to check the response body extremely often regardless of the status code being sent back.
If there's a way where we could say "this specific key will be this specific value in order to mean this specific scenario", I'd support that! But. Then. Isn't that the status that's I'm talking about?
I never said the error message can't or shouldn't be introspected and totally agree we should send our clients as much info as we can to make it clear on what they can do to fix any errors!
I just really really like being able to say "this specific field tells you, across all apis, what happened with this request" and don't see the gain in pushing that data into good luck finding out what key.
I'm currently dealing with an endpoint where it sends anywhere from 2-4mb worth of JSON data. Do I expect my clients to traverse every key and every deeply nested field to find what they're expecting for?
Absolutely not. because there's an internally standardized way of dealing with things and that is also precisely how we also deal with errors.
Now if you have an external API that faces many clients that you don't know about, then maybe there is consideration for usage of specific status codes.
But even then, your api should have documentation for it.
object graphs should be just called json or response body; you shouldn't use all these unorthodox terms.
a failure will be caught if a 400 was sent. This is how axios works - if you try a request, a 400 will automatically hit the catch.
a better designed system would be specificity. 75-90% of the time you're most likely going to look at the error body anyways.
I mean no personal offense but this is one of the most baffling and/or ignorant responses one could have imagined in this case.
json data => python server => redis
- json data gets serialized from json to string in the python server
- python server then stores that string in redis
to request from python server
- string is retrieved from redis server
- string is deserialized to send response as json
the Web client is in Javascript
JSON is native to Javascript
it is already readable to the client, there is no deserializing that needs to happen.
- It's non-standard, so different implementations will result in different formats
We have a standardized way of dealing with errors in which we will send 400/500
and have an error.data.err and an error.data.message in which err will give a title of the error, and the message elaborating the error in all of our errors across our APIs
the reason being is we don't need our developers to memorize dozens of generic status codes, and can be extremely specific in our error.data.err in which status codes can't.
It has worked for Facebook, It will work for us.
REST is now in practice an API using JSON over HTTP with appropriate verbs. And it's 'stateless' in that the server doesn't have to track the clients prior requests to understand the current request. That's about it I think....
Does anyone think that's stupid?
Agreed that everyone ignores the HATEOAS clause of REST, even when they call things "RESTful", but that's because the HATEOAS clause is tremendously stupid. I don't know why it's so common for narrow-minded intellectuals to decide that the reason the world is so complicated is because nobody has sat down yet and decided to make it simple. No, the world is complicated because reality has a surprising amount of detail [1][2]. When one person (or worse, a committee) sits down and decides "this is the way all of the world's information will be organized from now on" (HATEOAS, the "semantic web", etc.) they envision automated tools that can browse information the way Web Browsers browse the web. But as complicated as HTTP, HTML, CSS, JS, CORS, SVG, and now WebAssembly are, they are dozens of orders of magnitude simpler than "all the information anyone might ever want to make available over an API". Writing hypertext that can be understood by a tool to digest how your API should be consumed is just not possible. It doesn't work for people who are new to the API (they need to read the docs no matter what), and it's useless overhead for the people who use the API all the time (the hypertext is delivered with every request!?)
People don't use HATEOAS because HATEOAS is stupid.
[1] http://johnsalvatier.org/blog/2017/reality-has-a-surprising-...
I was a bit confused by that at first
From an implementation point of view, the difference between a POST, GET, and QUERY with a body is...trivial. I look forward to QUERY because it'll finally stop arguing about having to send a body on GET, or do a POST for something like a search.
It's gotten to the point where it all feels like naval gazing to me. I'd rather talk about what the api does rather than how amazingly restful it is.
I never understood why it became dogmatic though. Arguing for practicality over dogma seems like a losing battle, even in this industry where people are supposed to be relatively practical. I guess that's human nature :-/
Doing a POST on a /delete endpoint seem unnatural and I prefer serializing complex objects to JSON rather than to a query string.
I'm with you this far.
> From an implementation point of view, the difference between a POST, GET, and QUERY with a body is...trivial.
Sadly the difference is not trivial at all. I suppose if you're implementing an application-specific HTTP server that directly understands the application's semantics then the difference between these is trivial enough. But for the client the difference is not trivial at all, and the clients are bound to be off-the-shelf in many cases. And for an off-the-shelf HTTP server framework the difference may not be so trivial either.
GET /foo HTTP/1.1\rnHost: example.com\rn\rn{"foo":"bar"}
POST /foo HTTP/1.1\rnHost: example.com\rn\rn{"foo":"bar"}
QUERY /foo HTTP/1.1\rnHost: example.com\rn\rn{"foo":"bar"}
If it's a possibility that potential clients might refuse to do something like a GET with a body...don't use that as part of your api. Or make the endpoint respond identically to a GET with a body as a POST with a body. Similarly if you're working in a http server framework that makes such a thing difficult.At the end of the day, I care a hell of a lot more about how well the api is documented than I do about any of it's particular semantics.
And yes, I've seen proxies that don't support GET with a body.
> From an implementation point of view, the difference between a POST, GET, and QUERY with a body is...trivial
That said, I'm sure you're right, and there are proxies that don't support a GET with a body. It's an inadvisable thing to implement if there's a chance that your clients may not be able to support it. It's also really easy to just send a POST request with a body instead.
Where I get exhausted is when people start complaining that POSTing a body for something like a search query isn't restful. No, it's not, but it's not worth a heck of a lot of discussion either. In that respect, I'm happy that QUERY is being added, so it can hopefully resolve the distaste that the dogmatic experience. But that's really the only thing that makes me excited about it. Other than hushing the dogmatic, it doesn't really add much value.
> I get exhausted by people talking about integers, floats, and strings and arguing about the "types" of different things. From an implementation point of view, the difference between an integer, an float with or without a mantissa, or a string with numbers and decimal symbols is trivial. It has gotten to the point that...
Things have meanings, and in computer science and programming those meanings can mean a performace difference or a difference in results, so I think it is very important that we start to pay attention to the "naval gazing" aspects of these things and not use a GET with a body, since there is a QUERY, and other specs which might be kinda naval-gazey.
We are in a career that requires a pretty high level of strictness and we're always complaining about bugs in this or that software; then we argue that there is no meaning in certain things (only because in practice, people hadn't been strict about something like a GET request) which then leads to more bugs!
I certainly agree that things have meaning. They have such a meaning that we standardize them.
Thanks to the standardization and proliferation of the technology, there's generally an obvious correct choice, and a few poor arguments for others. In other words, it's generally not hard to make these things follow the Principle of Least Astonishment[0]. Every now and then there might be an interesting constraint that makes for a different choice (e.g., don't store zip code as an integer, because it may have a leading zero).
The amount of verbose discussion that I observe any time rest semantics come up still baffles me.
[0] https://en.wikipedia.org/wiki/Principle_of_least_astonishmen...
Some things have fundamental meaning. Like integers, floats, and strings are clearly very different things.
Whether you allow this or that via PUT vs. POST or whatever is just a convention, and one that you can freely change as long as it is fully documented.
I'm quite certain that if I were implementing a caching proxy (or a caching client, like a browser) - I'd find the diffence rather important. That said, I'm also sure I'd have to consider special exceptions for popular software that got the details wrong.
And if you're not caching, why use http anyway?
What did those poor ships do to you? /s
It's navel gazing, as in "looking at your belly-button": https://en.wikipedia.org/wiki/Navel
Thank you for the correction, I had a good chuckle.
Even without a cache, a GET and POST describe "what the api does" in detail - will my request change state on the server? Or will it be idempotent?. Considering them to be equivalent is like saying SQL statements SELECT, INSERT or DELETE are all the same. The distinction is absolutely critical.
edit: spelling
In other words, aside from <1% of developers, no one will care.
Also, graphQL seems to address many of the issues that manifest in the current form of GET. My bet will be that graphQL gets adopted more widely (and faster) than a new QUERY method.
I really don't like buying a[n] (HDMI|USB|Lightning) cable and then finding out it's not standards compliant, so my (PS5|PC|Phone) doesn't work.
I'm old enough to remember Web-development with Internet Explorer, and so old, I remember cross platform C++ development using Visual C++.
Nobody uses QUERY and it will thus have surprising consequences in no different amount than GET with a request. Firewalls will block you for "hacking". It's absolutely insane and idiotic that this is even a question. HTTP, which is a protocol that has no stated purpose, could have had no method at all and it would have made no difference. They could have just shoved whatever caching concern bullshit rationale which isn't sound either into more header lines. HTTP is a turing tarpit for protocols. Please close your websites, the users hate it and so do the engineers.
Clients of a service seldom get to meaningfully complain - use the service, don't use the service, whatever... So whatever the server wants, the server gets...
Although ironically I've also implemented servers and Gorton all kinds of complaining from clients - it's too hard to set a cookie, or header, can't we put tokens inside the JSON etc...
SOAP was the pinnacle of API design.
> When working with HTTP, there’s servers but also load balancers, proxies, browsers and other clients that all need to work together. The behavior isn’t just undefined server-side, a load balancer might choose to silently drop bodies or throw errors. There’s many real-world examples of this. fetch() for example will throw an error.
You have to separate HTTP, the spec, from REST the philosophy.
REST says you should use HTTP status codes. HTTP doesn't obligate you to use the full range of its status codes -- you can code up a resource that always returns 200 for POSTs, and that's fine as far as HTTP goes, and not as far as REST goes.
No, it doesn't, though leveraging the semantics of the underlying communications protocol (such as response codes when running over HTTP) is a convenient way to satisfying the “self-describing messages” constraint of REST.
> you can code up a resource that always returns 200 for POSTs, and that's fine as far as HTTP goes, and not as far as REST goes.
The only reason it would be even slightly problematic for REST is if it were inconsistent with the semantics of HTTP. Which, of course, returning 200 for anything but success would be.
I suppose a QUERY would really solve the problem. Maybe it's the best compromise.
Eventually I gave up and just ended the interview.
Last time I interviewed I pushed back but in a "are you sure? can we look it up?" kind of way and that seemed most palatable to them.
Both of these concerns are valid and need to be taken together. GET can be run just like a normal URL and thus is more versatile, but as you probably understand, less secure.
Where do I read about these features getting triggered?
If I'm interviewing someone who claims to be familiar with HTTP, I expect them to know this. They might choose otherwise in specific situations for specific reasons, but they should know the convention.
GET requests should be used to retrieve information, they're supposed to be read-only.
At most we could record like a counter or stats about the request, but not mutate resources internally.
The underlying resource (behind the cache) doesn't (shouldn't?) even know about the cache itself.
And the cache doesn't mutate any data points, just may store a version of it.
> GET: "means retrieve whatever information (in the form of an entity) is identified by the Request-URI" [2]
> POST: "used to request that the origin server accept the entity enclosed in the request as a new subordinate of the resource identified by the Request-URI" [3]
[1]https://datatracker.ietf.org/doc/html/rfc2616/
[2] https://datatracker.ietf.org/doc/html/rfc2616/#section-9.3
[3] https://datatracker.ietf.org/doc/html/rfc2616/#section-9.5