Node.js came out in 2009, a full ten years after HTTP/1.1 (RFC 2068) and its original http-parser is rather hard to follow, doesn't conform to the RFCs for performance reasons, and is considered unmaintainable by the author of it's replacement[0]
As for parsing HTML, well go look at how Cloudflare have stumbled[1]
[0] https://github.com/nodejs/llhttp
[1] https://blog.cloudflare.com/incident-report-on-memory-leak-c...
That's because of the way the parser is written. There are other simpler parsers that are much more readable.
[0] https://github.com/Samsung/http-parser/blob/master/http_pars...
[1] https://github.com/h2o/picohttpparser/blob/master/picohttppa...
i suspect the complexity you speak of is similar to MIME. where SMTP/POP/IMAP are pretty simple, things got pretty hairy with the introduction of MIME, SASL and friends.
i think, though, that most of the complicated stuff in http is optional, is it not? like if you don't send a header that compression is supported, the server won't compress... or am I misremembering?
either way, simpler to understand from a packet capture than a grpc stream or spdy/http2 stream.
The Joyent HTTP parser used by Node is very good but it's implemented in a way that makes the problem much more complicated than it needs to be. The biggest obstacle with high-performance HTTP message parsing is the case-insensitive string comparison of header field names. Some servers like thttpd do the naive thing and just use a long sequence of strcasecmp() statements. Joyent goes "fast" because it uses callbacks, which effectively punts the problem to the caller, and, for a few select headers which it handles itself, like Content-Length, it uses this really complicated internal "h_matching" thing for doing painstakingly written out hardcoded character compares. Redbean solves the problem by using better computer science: perfect hash tables. Thanks to gperf command. That makes the API itself much more elegant since the parser can not only go faster but return a hash-table like structure where individual headers can be indexed without performing string comparisons.
I think there are already some versions of the ragel code online, but they might be for other target programming languages.
Based on this consideration I was thinking that the ragel state machine would generate faster code for the non happy path (invalid non-ascii, or other types of error) at least in the GOTO version.
When working on the full list it makes perfect sense to check the minimum amount of bytes for identifying headers, so thank you for the clarifications, very informative. :)
The charset problems alone are a nightmare.
Parsing the wire format is pretty breezy, (Don't forget trailers!)
See the remark at the bottom of the Boost Beast parser docs for a hint at the trade off here:
https://www.boost.org/doc/libs/1_75_0/libs/beast/doc/html/be...
I think the implementation is gold standard:
https://github.com/boostorg/beast/blob/develop/include/boost...
> Line folding is forbidden.
Okay, I'll bite: What are you talking about? Fastest under what conditions, using what measurements? What is the project, and where is the analysis? EDIT: you probably mean redbean (https://redbean.dev/) - the source of which is, oddly, embedded within cosmopolitan. However the question about analysis stands.
I really like it and will attempt to use it for something real. However, I fear that even this bit of magic doesn't address the central problem of our time, which is software distribution. I believe that the web has solved that problem, and although the web is currently abused by central power, and webapps tend to be thin, animated protocol viewers, it doesn't have to be that way. You've created/discovered a local (maybe global, given real-world limits) minima of what a binary executable can be, but this only finds the minima of the pain of traditional software distribution, but doesn't eliminate it.
The real path forward, if I might be so bold, is to make a browser on top of cosmopolitan/redbean, and bring TBL's original dream of a singular client+server http/html runtime to modern fruition - but with additional superpowers that cosmo brings which I don't think TBL anticipated. No doubt some enterprising souls are already working to get Bellard's QuickJS into redbean to mimic node. Then you need window/drawing context, and the rest of the browser, including layout, could be done in (presumably equally tight) JS. Have you given any thought to exposing those drawing syscalls directly instead of delegating to the browser? And if you haven't and are interested, may I suggest Java's AWT v. SWT as an interesting case study in "where the indirection should go".
For example, Content-Length isn't just a single header with an integer, like the spec says. You need to support responses with multiple Content-Length headers and comma-separated lists of potentially contradictory lengths, and then perform garbage error-recovery the way Internet Explorer or Chrome did. Getting this wrong will make your client hang, consume garbage, or allow response stuffing.
https://github.com/web-platform-tests/wpt/pull/10548/files
This problem does not exist in HTTP/2.
Multiple responses, especially with chunked encodings, are really really hard to get right. Even more so when you have to be able to resume downloads due to socket instability.
I drop two things here to counterargue that HTTP is easy to parse: 206 multiple ranges (requests will not always be responded to, and ranges from servers are almost always invalid when resumed) and Transfer Encodings (br inside gzip inside deflate inside chunked, anyone?)
Both headers are implemented spec incompliant by every single web server I've seen, including nginx, apache and others.
Source: building a peer to peer web browser that shares its bandwidth (and cache and downloads) with trusted peers. [1]
You should not expect the Node.js parser to be simple.
I've worked heavily on some HTTP implementations and its ridiculously hard to get them right.
Not to mention, this "server" only responds to a simple well formed GET request. Without handling about 90% of what the HTTP specifications talk about. Its a nice project, but it doesn't speak to the simplicity of HTTP
> this "server" only responds to a simple well formed GET request.
And not even that. The Request-URI in a Simple-Request line (inherited from HTTP/0.9) may contain escape characters. (e.g. `GET /my%20file.txt` to get `my file.txt`) HTTP/1.0 states "The origin server must decode the Request-URI in order to properly interpret the request."[1] This server does not.
Which is not to say that this server isn't interesting. Just that it's not a demonstration of how easy HTTP/1 is to parse.
[1]: https://www.w3.org/Protocols/HTTP/1.0/spec.html#Request-URI