What every web developer must know about URL encoding
blog.lunatech.com
blog.lunatech.com
The reserved characters are different for each part
But if it's true that you can _optionally_ percent-encode any character in any part of the URL (path or query) -- is this true? I think so -- then there is in fact a way to _encode_ any URI without as much syntactical awareness of the URI structure.
This means that the "blue+light blue" string has to be encoded differently in the path and query parts: "http://example.com/blue+light%20blue?blue%2Blight+blue". From there you can deduce that encoding a fully constructed URL is impossible without a syntactical awareness of the URL structure.
Does it? Could you optionally encode it as:
http://example.com/blue%2Blight%20blue?blue%2Blight%20blue
That is, percent encode (not the legacy encode-space-as-plus) both the plus and the space in both path and query?* You'd still need enough syntactic awareness of the URI structure to know to leave path-seperator "/" alone, and query-part-beginning "?" alone.
* And you'd need even more syntactic awareness to properly _decode_ URIs, which may not have chosen to optionally always-percent-encode.
* But it might be wise if libraries chose to encode things in that consistent way, for instance encoding plus in path even though it's not required, never encoding space as plus even in query.
* I think there is _no_ good reason for any modern library to encode spaces, even in the query string, as "+" instead of %20. Even though many of them do. Don't do it.
* I don't use Java much, and couldn't entirely follow those examples (this stuff is confusing to talk about -- as in all escaping issues) -- but it sure does seem like those parts of stdlib/commonly used libraries are pretty darn broken. (And I'm sure Java is not alone here -- there is a long history of devs being confused about this stuff. Again, as with just about any escaping-related issue).
I mean, who thought it was a good idea to have different escaping for different parts of a URL? That's like having a car that uses a lever to turn right and a foot pedal to turn left.
Of course, the answer is: nobody. URLs weren't designed, they grew organically. The result is the ubiquitous mess we have now.
It does work. The organic, evolving web has survived and thrived while sanely-designed standards have withered and died. Maybe it's ultimately the best way to do things.
But it still bugs me that I have to deal with such weird, awful nonsense when I need to do something related to the web. TCP/IP, while certainly not lacking in quirks, is still a million times better. It's sad that we couldn't end up with something a little more consistent at the higher levels.
HTTP and SIP, for instance, have ridiculous parsing rules and tons of edge cases. The SIP authors even put together a "torture test" RFC where they take glee in making the most insane messages that still parse. And when they don't parse, they suggest the implementation infer the meaning. Seriously.
HTTP and SIP allow newlines in headers that get consumed in parsing, so you can manually word-wrap lines. They allow comments in HTTP messages. SIP (HTTP too?) allows headers _in the URL_.
When a non-committee (TCPIP) or non-academic (SPDY) entity does a spec, they tend to remove a lot of this cruft and realise that human-readability comes second (unlike in programming languages).
Assuming Markdown, that would be::
[Java](http://en.wikipedia.org/wiki/Java_(programming_language))
Bad parser.There's also some difficulty with how RFC 2616 (the current HTTP/1.1 spec) demands using definitions of URLs from RFC 2396 which is not the current URI spec, but the update URI spec seemed to me to be inconsistent with general usage of HTTP. If I remember correctly, it ended up implying that you cannot have a query string without sending your scheme as well, but I may be wrong.
The URI spec is hideously complex, but also very comprehensively defined in those RFCs. It's an interesting job to look through them.
| you cannot have a query string without sending
| your scheme as well, but I may be wrong
I don't remember reading that in the RFC.Also, I'm curious for examples of URLs that break RFC 3986[1].
e.g. a link to
http://example.com/?colour=blue&age=old
with link text balloon
should appear in html as <a href="http://example.com/?colour=blue&age=old">balloon</a>
but might erroneously appear as <a href="http://example.com/?colour=blue&amp;age=old">balloon</a>
...or if a (simple) client neglected to html-decode the uri for some reason, then the link would still work.Whether this was officially recommended or was merely a folk recommendation, I don't know.
Does not work: http://www.foobar.com/api?v=2/get?item=1
Does work: http://www.foobar.com/api;v=2/get?item=1Of course, it does depend on what parser gets them first!
Maybe I'm jaded. As a newly-minted SDET I once tried to test email addresses as defined by RFCs and eventually realized that I was trying to test adherence to standards instead of what actually happened.
My humble suggestion is to use a well-vetted URI-handling library and employ a combination of eyeballing it and some faith. (I hate to recommend faith but there you go.) And if there's not already a lib for that in Java (gotta be but I try to avoid Java) maybe some Java-head could write one for everyone . . .
Shame, I can see a use for them.
basically, unlike Java, it doesn't give you an encode() function that takes an arbitrary string... the only urlencode() function expects data representing a query
obviously, you still have to remember to handle the quoting of each part of the url separately... if you build your url (actually just resource+query+fragment) , and then you just quote() it at the end, you're no better than with java
e.g, if you have a path made by 2 segments "yadda/yadda" and "foo/bar"
quote("/".join(["yadda/yadda", "foo/bar"]))
yields 'yadda/yadda/foo/bar', which might not be correct, if what you want is actually
"/".join(quote(segment, safe="") for segment in ["yadda/yadda", "foo/bar"])
that yields 'yadda%2Fyadda/foo%2Fbar'
kinda error-prone, if you ask me
Also, python's urlparse seems to not handle correctly path parameters:
urlparse("http://example.com/egypt;p=0/nile;p2=1;p3=2")
only recognizes p2=1;p3=2 as path parameters
I think that part of the confusion is that we think of encoding as "The way in which symbols are mapped onto bytes", but if we use that meaning, it's not correct to talk about "url encoding", because each part of the url cannot be converted in ascii while ignoring the context (are the / meaningful?) and the place it appears into the url
it's more like a "url language", and if we would talk about "parsing" or "formatting" imho we could get less ambiguities and misunderstandings
PS: I realized just now that for the "http://example.com/egypt;p=0/nile;p2=1;p3=2" example, the doc suggests to use the urlsplit function instead (that will avoid to parse the path parameters altogether... kind of a non-solution imo, but at least it's known)
Guava also has tonnes of good stuff - http://docs.guava-libraries.googlecode.com/git/javadoc/com/g... , and its ostensibly lighter weight and more modular too.
Sadly true. Drovr me nuts trying to analyze a URL on Apache/PHP.