HTTPWTF
httptoolkit.tech
httptoolkit.tech
Chunk extensions. Most people know HTTP/1.1 can return a "chunked" response body: it breaks the body up into chunks, so that we can send a response whose length we don't know in advance, but also it allows us to keep the connection open after we're done. What most people don't know is that chunks can carry key-value metadata. The spec technically requires an implementation to at least parse them, though I think it is permitted to ignore them. I've never seen anything ever use this, and I hope that never changes. They're gone in HTTP/2. (So, also, if you thought HTTP/2 was backwards compatible: not technically!)
The "Authorization" header: like "Referer", this header is misspelled. (It should be spelled, "Authentication".) Same applies to "401 Unauthorized", which really ought to be "401 Unauthenticated". ("Unauthorized" is "403 Forbidden", or sometimes "404 Not Found".)
Also, header values. They're basically require implementing a custom string type to handle correctly; they're a baroque mix of characters & "opaque octets".
Plus surely there are many crusty middleboxes that will break if anybody tried to use that feature. Remember all the hoops websockets had to jump through to have much of a chance working for most people because of those? Many break badly if anything they were not programmed to handle tries to pass through.
Oof, I hadn't mentally connected those dots, but you're completely right. (As Transfer-Encoding is hop-by-hop, not end-to-end…)
You are supposed to treat all of them as "opaque octets"... or something like this might happen:
header_as_raw_bytes == b"chunked"
which I would argue is still decoding the header: your language of choice had to encode that string into bytes in some encoding in the first place, so even though you're comparing the encoded forms, there's still a character encoding at work.But, some of the headers are case-insensitive. E.g., Content-Type, Accept, Expect, etc.
That golang bug is precisely not treating the non-characters (the "opaque octets", as defined by the standard, that is, the octets that form obs-text) as if they were characters. You won't hit that bug in the safe subset, presuming you're implementing other parts of the standard correctly. (Which is… a huge assumption, given HTTP's complexity, but that's sort of the point here.)
Lots of servers I've encountered with my browser stealth actually violate the spec and send "lengths" that do not match the sent payload lengths afterwards. Some reverse proxies also mess up the last chunk, so they're violating the spec there, too, and send a chunk with a negative length...and the spec doesn't even define how to handle this. I've also seen servers send random lengths in between, but without a payload that follows.
I would also like to add range requests (206 partial content) here. In practice, it's totally unpredictable how a server behaves when requesting multiple content ranges. Some reply with no range at all, even with correct headers. Some reply with more or less ranges than requested. Some even reply with out of boundary ranges that are larger than the content length header of the same response because they seem to use a faulty regexp on the server side.
It's a total shitshow.
Always wondered if developers found that easy.
I just strip out the chunk lengths with a filter, suitable for use in UNIX pipes. It's like three lines in flex. I have always been aware of the different things that servers could "legally" do with chunking from reading the HTTP/1.1 spec but as the parent says no ever does anything beyond the basic chunk lengths. For example, how many servers support chunked uploads.
With the filter I wrote, as crude as it is, I have never had any problems. Works great with HTTP/1.1-pipelined DoH responses.
Hmm. I suppose it isn't explicitly called out, but I think it's fair to say that such a request is a 400 Bad Request, as it doesn't match the grammar. (There's no possibility for a negative chunk length, as there's no way to indicate it.)
What you pass in the "Authorization" header is an user identity, which is established through authentication. And the server uses this identity to decide if you are authorized.
You can't actually know whether the outgoing side of a TCP socket is closed, unless you write something to it. But it's hard to come up with something to write to an HTTP/1.1-over-TCP socket before you respond with anything, that would be a valid NOP according to all the protocol layers in play. (TCP keepalives would be perfect for this... if routers didn't silently drop them.)
But I guess sending an HTTP 102 every second or two could be used for exactly this: prodding the socket with something that middleboxes will be sure to pass back to the client.
If so, that's awesome! ...and also something I wish could be handled for me automatically by web frameworks, because getting that working sounds kind of ridiculous :)
Wouldn't setting appropriate net.ipv4.tcp_keepalive_* and trying to read work?
Like I said:
> TCP keepalives would be perfect for this... if routers didn't silently drop them.
There are lots of middleboxes that don't pass along empty TCP packets. TCP keepalive is in a similar situation to IPsec: great for an Intranet, or for two public-Internet static peers with a clear layer-3 path between them; but everything falls apart in B2C scenarios.
Plus, to add to this problem: HTTP has gateways (proxies et al.) Doing TCP keepalive on the server end, only tells you whether the last gateway in the chain before the server is still connected to the server, rather than whether the client is still connected to the server.
Unless you can get every gateway in the chain to "propagate" keepalive (i.e. to push keepalive down to its client connection, iff the server pushes keepalive down onto it), silent undetected TCP disconnections will still happen—and even worse, you'll have false confidence that they aren't happening, as all your sockets will look like they're actively alive.
For what I'm doing, the client end isn't likely to have any gateways, so TCP keepalives "would be" workable for my use-case if not for the middlebox thing. But in full generality, TCP keepalives aren't workable, because there's always those corporate L7 caching proxies + outbound WAFs messing things up, even when L4 middleboxes aren't.
Keep your TCP keepalives for running connection-oriented stream protocols within your VPC. For HTTP on the open web, they're pretty unsuited. You need L7 keepalives. (If you've ever wondered, this is why websockets have their own L7 keepalives, a.k.a. "ping and pong" frames.)
> and trying to read
An HTTP client connection can legally half-close (i.e. close the output end) when it's done sending its last request; and this will result in a read(2) on the server's socket returning EOF. But this doesn't mean that the client's input end is closed! You have to do a write(2) to the server's socket to detect that.
And, since empty TCP packets aren't guaranteed to make the trip, that means you need to write a nonzero number of bytes of ...something. Without that actually messing up the state-machine of your L7 protocol.
Golang's GRPC library implements keepalives at the GRPC protocol level [1]. It provides a `Context` value [2] that code can use to detect peer disconnect and cancel expensive operations.
Golang's HTTP server API does not provide any way to detect peer disconnect before sending the final response [3].
Rust cannot set SO_KEEPALIVE [4]. One could possibly implement keepalives by writing zero-length chunks to the socket.
Java's Netty server library can set SO_KEEPALIVE [5]. One can then code a request handler that periodically checks if the socket is connected [6] and cancels expensive operations. Unfortunately, there is standard tooling to do this.
EDIT: You did mention TCP keepalives. I was not aware that some routers drop them. Can you link to any data on the prevalence of tcp-keepalive dropping for various kinds of client connections: home router, corporate wifi, mobile carrier-grade-NAT?
[0] https://tldp.org/HOWTO/TCP-Keepalive-HOWTO/overview.html
[1] https://pkg.go.dev/google.golang.org/grpc/keepalive
[2] https://pkg.go.dev/google.golang.org/grpc#ServerStream
[3] https://pkg.go.dev/net/http#HandlerFunc
[4] https://github.com/rust-lang/rust/issues/69774
[5] https://netty.io/4.1/api/io/netty/channel/ChannelOption.html...
[6] https://netty.io/4.1/api/io/netty/channel/Channel.html#isOpe...
I can point out the obvious "analytical evidence", though: note how all the platform APIs that did expose SO_KEEPALIVE are from the 90s at the latest — i.e. before the proliferation of L4 middleboxes. And note how modern protocols like Websockets, gRPC, and even HTTP/2 (https://webconcepts.info/concepts/http2-frame-type/0x6) always do their own L7 keepalives, rather than relying on TCP keepalives — even when there's no technical obstacle to relying on TCP keepalives.
Iirc, the context mechanism can be used to detect the client's disconnection or cancellation of the request in some cases. From [1]:
> For incoming server requests, the context is canceled when the client's connection closes, the request is canceled (with HTTP/2), or when the ServeHTTP method returns.
The TLS 1.3 spec states "Zero-length fragments of Application Data MAY be sent, as they are potentially useful as a traffic analysis countermeasure."
I guess that tls libraries wouldn't expose an api to do that which is problematic for this approach.
This problem is one reason why success/error should NOT be the first thing to send. It should be a trailer.
(HTTP/HTML tendency to substitute the response body for a human-visible error would require another mechanism to "reset" the response body.)
There is no current browser, or client that by default will send an Expect: 100-Continue.
cURL removed it because it was too often broken. See https://curl.se/mail/lib-2017-07/0013.html
As of right now, while server authors will continue to need to support it, it is unlikely that it is a well tested code path, and it will likely break in weird ways even trying to use it.
So pre-approval for a large file upload is not even valid anymore.
Best to just never use custom headers.
I've written more about this here: https://developer.akamai.com/blog/2015/08/17/solving-options...
That's why they're so great. use a custom header and never worry about CSRF issues.
Use custom header and be sure that if request comes from the browser it was made by legitimate code from your origin.
This becomes useful though if you send a request including a Except: 100-continue header. That header tells the server you expect a 100 response, and you're not going to send the full request body until you receive it.
I’m guessing that should be Expect?
Overall interesting article, thanks for writing it!
Nginx gives you a 400 Bad Request response, Apache does nothing, and other servers vary in whether they return a non-error code.
HTTP/1.0 400 Invalid HTTP Request
My mistake, or are there other working end-points out there (I tried google, yahoo, and cbc.ca)?You can include the same header multiple time in a HTTP message, and this is equivalent to having one such header with a comma-separated list of values.
Then there's WWW-Authenticate (the one telling you to re-try with credentials). It has a comma-separated list of parameters.
The combination of those two leads to brokenness, like how recently an API thing would not get Firefox to ask for username and password, because it happened to have put "Bearer" before "Basic" in the list.
The Set-Cookie header (sent by the server) should always be sent as multiple headers, not comma separated as user agents may follow Netscape's original spec.
https://stackoverflow.com/questions/2880047/is-it-possible-to-set-more-than-one-cookie-with-a-single-set-cookie
https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Set-Cookie
On the other hand in HTTP/1.1 the Cookie header should always be sent as a single header, not multiple. In HTTP/2, they may be sent as separate headers to improve compression. :) https://stackoverflow.com/questions/16305814/are-multiple-cookie-headers-allowed-in-an-http-request[1]: https://fasterthanli.me/articles/aiming-for-correctness-with...
https://www.w3.org/Protocols/HTTP/HTRQ_Headers.html
This is not a custom X- header but an official header. Also Email: and some other odd headers were standardized at that time.
That's why you spelled 'spelled': spelt :D
(Edit: or perhaps it's more about active/passive voice? Thinking particularly about burnt/burned.)
It's taken to an extreme by some thicker accented (dialected?) people around where I grew up though - an' so I turnt [turned] round (right-right round) an' said to 'im [...]!
UK native, upon seeing the word "color" thinks: "Aha, US spelling!"
US native, upon seeing the word "colour" thinks: "Aha, a typo!"
- readily available spell checkers
- the international interplay of ideas
The former stabilizes spelling, the latter says "screw it, through is spelled thru, I made up yeet, and emoji are valid grammar". But also grammar and spelling Nazis.
My money is on fluidity of language.
Could someone explain why this needs a new status code at all? At the point where the new status code sends "early headers", the client was expecting the regular status code and headers anyway. Why could the server not simply do:
1) Receive request
2) Send 200 OK and early headers, but only send a single trailing newline (i.e., terminate the status line and last early header field, but don't terminate the header list as a whole)
3) Do the actual request processing, heavy lifting, etc
4) Send remaining headers, double-newline and response body, if any.
On the client side, a client could simply start to preload link headers as soon as it receives them, without waiting for the whole response.
This seems like it would lead to pretty much the same latency characteristics without needing to extend the protocol.
The only major new ability I see is to send headers before the (final) status code. But what would be the use-case for that?
Edit:
The RFC[1] sheds some light on this: The point seems to be that the headers sent in an 103 are only "canon" if they are repeated in the final response. So a server could send a link header as an early hint, then effectively say "whoops, disregard that, I changed my mind" by not sending the header again in the final response.
I still don't see a lot of ways a client could meaningfully respond to that, but I guess it could at least abort preloading to save bandwidth or purge the resource from the cache if it was already preloaded.
Fun fact, reddit used to have 'X-Bender: Bite my shiny metal ass' on every response. Sadly they seem to have removed it.
> Disallow: /my_shiny_metal_ass
This was there when I checked just now; was it removed and re-added?
You thought you were replying a level higher.
The Rust Book has an awesome "final project" where it walks you through building a multi-threaded web server. If you're a battle-hardened C/C++ dev looking for an inroad to Rust, this is a great place to start.
https://tools.ietf.org/html/rfc5987
E.g. sending the file "naïve.txt" using the Content-Disposition header. Content-Disposition: attachment; filename=na_ve.txt; filename*=utf8''na%C3%AFve.txt
https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Co... The parameters filename and filename* differ only in that filename* uses the encoding defined in RFC 5987. When both filename and filename* are present in a single header field value, filename* is preferred over filename when both are understood.This made me laugh.
Incidentally, reminds me of a company I used to work at where one of the devs thought it was hilarious to return 418 (I'm a teapot) [1] for all bad requests. Unfortunately sometimes these were actually 5xx-level errors so it quickly become annoying. The Twilio header listed above seems fairly innocuous though.
Do try out HTTP Toolkit and let me know what you think, but it's not a general purpose HTTP client like Postman or HTTPie. It's actually an HTTP debugger, more like Fiddler/Charles/mitmproxy, for debugging & testing. A convenient HTTP client is definitely planned as part of that eventually, but not today.
Browsers store, proxies cache , so it should be no-cache, obviously!
Sure, it's stupid, but naming is hard and these things happen all the time.
The "no-cache" was a hint not about caching the content but about caching subsequent requests, and could optionally specify specific fields that would indicate that a new candidate request needed to be sent to the server as the content might be different. There's this reality that just to render content, the browser effectively must have a cached copy of the content, so the notion that the response wouldn't be cached wasn't really even in the cards. Whether you used the cache or not was a decision made at the time you were sending a request, not when you were consuming the response.
The "no-cache" directive meant, "hey, don't check for a cached copy of the content, just go fetch new content". It was often used by analytics pieces so that the server could count how often content was looked at.
Back in the day you had terrible latencies (particularly over dialup). You also had issues with horribly asymmetric bandwidth that meant the data you sent could become the bandwidth bottleneck (outbound bandwidth constraints would mean ACK packets would get queued up, delaying downloads even when you had plenty of download bandwidth), and of course HTTP requests weren't terribly compact, so this could really make a big difference.
Caching requests was a big deal. Performance could be improved significantly by "cheating" and just not sending a new request, and this lead to some very aggressive caching strategies. The, "check if the content really is different, and just use the original copy if they aren't" hack a pretty common one. If nothing else, it saved the browser the overhead of re-rendering the page and the accompanying annoying user experience of seeing the re-render.
The original protocol didn't have any notion of no-store, and specifically mentioned that "private" didn't really provide privacy, but more that the content should be "private" in the sense that only the browser itself should store the content. Again, there's an assumption that the browser is going to put everything it gets into a "cache", because it has to.
You could use "max-age", but a lot of caches would still shove the object in their cache and only expire it on a FIFO basis or when a new request was to be sent (and it was vulnerable to clock skew problems). Sounds dumb, but it was the kind of dumb that kept code simple and worked pretty well.
So now that the practices were in place, you need a new directive to say, "hold up, that old approach is NOT a good idea here". So they came up with "no-store" as a way to say, "don't even put it in the cache in the first place".
That might be a little counter-intuitive, but if you read the definitions of the words, it does make sense.
> no-cache won't give you cached entry without validation
> must-revalidate, max-age=10 will revalidate only after the time has expired
must-revalidate: you MUST revalidate with the server before using this after it expires
They overlap but address different things; no-cache is for cacheability while must-revalidate is for validity. I don't think either of them are named very well.
Private? No. Cache!
https://beta.shodan.io/search/report?query=http.status%3A%3E...
I'll enjoy knowing this next time I reread Going Postal :)
add_header X-Clacks-Overhead "GNU Terry Pratchett";
in your server{} block.This is popular and supported because it serves a real need. A GET request shoves all the query parameters into the URL, and if that gets long enough, it will be truncated somewhere and the results will be wrong or you'll just get an error.
By instead sending a POST request, where all the parameters go in the body and don't get truncated, but using the header to tell the server to treat it as a GET request, you solve those bugs almost seamlessly.
The drawback is that you don't have a bookmarkable/linkable URL - some of the semantics are lost. But that can be worked around afterword with an id or hash of a previous request's parameters that tells the server "give me the same results as if I had entered all the query parameters that were entered in this previous request".
It's not optimal, but pragmatic.
The origin headed should be the only way to fix these issues which would force web frameworks and servers to check this header by default. Way simpler, authentication (authenticating the requestor site, not the end user) is done on the server not the client, errors can actually be handled since don't fail silently, and failures appear in server logs as well as the client.
Does amazon.com really make it's page more performant by sending 25 chunks less than 2k, some less than 50 bytes, while I'm trying to grab 115k for a page?
It's all so weird to me.
Those 2kB is a bit too small for top network performance, so you may see a negative impact. But if they increase it to something like 10kB, it's harmless.
You are right that it's not a great thing to do. A little bit of buffering on the sender can improve things a lot. But it's an easy thing to do, so people do it.
This works even better in HTTP/2 or HTTP/3 / QUIC. A Go server reading from a lot of microservices can produce pretty weird output on HTTP/2 because now not only is it in odd sizes determined by network timing, it doesn't even need to be in order.
What does the author mean by this? Why can't a "resource request" include custom headers? I am assuming that a "resource request" is just a non AJAX request. Any HTTP client should be able to include whatever headers they want no matter the source.
I was always a fan of setting this to either be silly ( X-Powered-By: magical elves ) or outright lie (Tell em it's some ruby thing when it's ASPNET)
But also...cool tool!
When things happen on the site and it's shown to the customer immediately via WS, it's just a delightful experience.
Unfortunately the browser javascript "fetch API" still doesn't implement streaming request bodies. If it did, websockets would be obsolete.
I once wrote a prototype video surveillance system with that, sending multiple still images, before video streaming was a thing. 1997, that was a long time ago...
JP!
I figured after a week of reading comments from so many beautiful brains who better to ask ya know? I wish I would not have gotten in a wild group to not get beat up all the time when young and stayed on the path I was on. Those beatings are nothing and I could deal with them all every ay and be better off than my current day to day.
MAYBE\ All I know is I know nothing at all for certain.
I am grat t business and cannot even get funds enough together for any business unless I wanted to just Break Bad ( I certainly could now that I have met so many people society would think are terrible, but they are not and offer to help if I ever can blah blah blah...
THANK YOU very much for reading.