Parsing millions of URLs per Second (2023)
onlinelibrary.wiley.com
onlinelibrary.wiley.com
Why would normalization change http:// to https:// ?
OTOH I'm browsing since years forcing HTTPS only and life goes on fine. If the absolute worse comes to worse, I can use archive.is or archive.org but it's very rare that I need that.
Basically: if a link is HTTP to me it's not worth opening.
The one exception would be Debian packages URLs: but these are signed and the signatures are verified.
User _apt is the only one allowed to emit HTTP traffic.
This prevents my ISP or anyone else injecting nasty stuff.
> For example, given the base string http://example.org/foo/bar, the relative string http://example.com/ leads to the final URL http://example.org/example.com/.
That’s just… no. I do not believe I have ever encountered any software which would parse it in that way, and I refuse to believe such software ever existed. It would be <http://example.com/>.
But the PDF matches the HTML. I dunno, something weird is going on. Look at the hyperlinks there, too, “http://xn--ivg but not the rest of the URL that follows, and how the -- has been changed to –. Something went wrong somewhere in the editing or publication.
I dislike automatic linkifiers, especially in technical contexts, because they get things wrong so often, as regards what is a link at all (and certainly never linkify if there’s no protocol! “example.com/foo” should not be turned into <http://example.com/foo>), and as regards what can be part of the link (largely around trailing punctuation). Just require explicit delimition, like <…>, or else it’s text.
(Markdown’s […](…) is bad because ) is part of URL code points, meaning parentheses in URLs won’t be percent-encoded by a normal serialiser, so then its parser gets messy trying to compensate, assuming that parentheses will normally be paired in URLs. Your delimiter needs to not be part of the set of URL code points.)
HN’s auto-linkifier is, most of the time, one of the better ones (it was bad ten years ago, but got fixed around punctuation inclusion a few years ago), but it still has problems. I noticed too late that it mangled something in my comment: where you get http://xn--ivg, that xn--ivg is ”, because what I actually wrote was
… too, “http://” but not …Are the benchmarks comparing node versions valid to conclude a real world performance increase?
one possible confounder is the version of V8.
https://github.com/nodejs/node/blob/v18.x/deps/v8/include/v8... https://github.com/nodejs/node/blob/v20.x/deps/v8/include/v8...
ideally, they would've patched Node 18.15 with their changes directly and test their patch against 18.15.
https://daniel.haxx.se/blog/2023/11/21/url-parser-performanc...
> Parsing millions of URLs per second
I caught that one manually but YOShInOn's tail end needs some love and could be updated so it that it fixes up titles that get mashed automatically or adds a comment sometimes to editorialize or provide an archive link.
A pdf link to "5 Reasons To Do Things" will be "Reasons To Do Things [pdf]" for example.