In search of the perfect URL validation regex
mathiasbynens.be
mathiasbynens.be
I should mention that 99.9% of domains will fall into standard form ( handle.domain or ip.ip.ip.ip )
As such, You are definitely more likely to let a user enter a bad URL they did not intend because it validates then to let a uncommon domain actually be used.
As such- a much simpler regex would likely 'make more people happy' than being 100% correct to tech spec.
Or, you know, use a robust existing validator or parser. Like PHP's, for instance.
[1] https://github.com/TazeTSchnitzel/Faucet-HTTP-Extension - granted, this deliberately limits the space of URLs it can parse, but it's not difficult to cover all valid cases if you need to
[2] https://github.com/php/php-src/commit/36b88d77f2a9d0ac74692a...
I also met very few people that actually understood regex. It's a whole new skill and if I use it in my application I don't know if the next guy can pick it up.
We should get a port to every language.
[1]: http://stackoverflow.com/questions/1732348/regex-match-open-...
(I still think it's not a great idea. Being regular isn't necessarily the same as being parseable with a maintainable regex.)
I actually have a use-case. I am firming up a feature right now that detects when a user types a url into a text field and replaces it on the fly with a footnote-style reference number (much like your comment above). This is done to (1) minimize input string length, (2) draw the benefits of a consistent interface, and (3) avoid screwing around with the fragility of url shortening nonsense.
I may regret this, but here's a link to my dev environment for this feature (please be gentle), to see it in action:
https://cloudcity.tenfourgood.com/cloudcity
...just start typing in the big text box and add in some urls.
It uses a fairly ugly-looking regular expression[0].
If you take out the unicode mumbojumbo, it's not really THAT tricky of a pattern. It does fail on IP addresses and it may be a little over-aggressive on matching, but, I wanted it to catch things like "abc.com" and "//xyz.com".
Edit: Formatting and clarity. Removed the explicit regular expression because it predictably got garbled.
[0] http://regexr.com/38vsq not exactly the same one I'm currently using, but it's pretty close. See source for most up-to-date version.
The regular expressions presented do not conform to the URL Standard or to any RFC, but rather to the list of requirements on that page.
I applaud your work in improving PHP’s built-in URL parser.
The @stephenhay is just about perfect despite being the shortest. The subtleties of hyphen placement aren't very important, and this is a dumb place to filter out private IP addresses when a domain could always resolve to one. Checking if an IP is valid should be a later step.
"b--c" seems to be a valid domainlabel, though, so I'm not sure why it's on there.
RFC2396§3.2.2 defines hostname as:
hostname = *( domainlabel "." ) toplabel [ "." ]
RFC3986 which obsoletes RFC2396, seems to make fewer claims about the authority section, and specifically says in Appendix D.2 that the toplabel rule has been removed.I disagree. But, this is a tricky one. The relevant specs are:
Spec | | Validity | Definition
URL (RFC1738) | obsolete | invalid | hostname = *[ domainlabel "." ] toplabel
HTTP/1.0 (RFC1945) | current | invalid | host = <A legal Internet host domain name
| | | or IP address (in dotted-decimal form),
| | | as defined by Section 2.1 of RFC 1123>
HTTP/1.1 (RFC2068) | obsolete | invalid | ; same as RFC1945
HTTP/1.1 (RFC2616) | obsolete | valid | hostname = *( domainlabel "." ) toplabel [ "." ]
URI (RFC3986) | current | valid | host = IP-literal / IPv4address / reg-name
| | | reg-name = *( unreserved / pct-encoded / sub-delims )
HTTP/1.1 (RFC7230) | current | valid | uri-host = <host, see [RFC3986], Section 3.2.2>
The only way that URL is invalid is if we are in a strict HTTP/1.0 context.As a note about RFC1738 being obsolete: these days a URL is just a URI (1) whose scheme specifies it as a URL scheme, and (2) is valid according to the scheme specification.
As the given URL is a valid URI, and is valid according to the current http URL scheme specification (RFC7230), that URL is valid.
The goal was to come up with a good regular expression to validate URLs as user input, and not to match any URL that browsers can handle (as per the URL Standard).
a.b--c.de is supposed to fail because `--` can only occur in Punycoded domain name labels, and those can only start with `xn--` (not `b--`).
Trying to shoehorn NFAs into parsing stuff that isn't a regular expression is generally a bad idea. (See: Langsec.)
If you really want your URL shortener to reject bad URLs, then you need to actually test fetching each URL (and even then...)
As an aside, I'd instantly fail any library that validates against a list of known TLDs. That was a bad idea when people were doing it a decade ago. It's completely impractical now.
It's useful to find and linkify URLs in text (e.g. in your HN comments, how do you think HN makes http://foo.com into a link?)
In that scenario, and with the given requirements, I can’t think of a case where the validation fails. There’s no need to worry about protocol-relative URLs, etc.
(Keep in mind that this page is 4 years old — I very well may have missed something.)
> If you really want your URL shortener to reject bad URLs, then you need to actually test fetching each URL (and even then...)
I disagree. http://example.com/ might experience downtime at some point in time, but that doesn’t mean it’s suddenly an invalid URL.
> As an aside, I'd instantly fail any library that validates against a list of known TLDs. That was a bad idea when people were doing it a decade ago. It's completely impractical now.
Agreed.
There are some examples of these pathological inputs at https://github.com/tornadoweb/tornado/blob/master/tornado/te...
http://swtch.com/~rsc/regexp/regexp1.html
Because the issue with the URL regex mentioned is with backtracking.
https://github.com/PiPeep/NotVeryCleverBot/blob/coffee-rewri...
Note the commented out lines in the here-regex.
uri = URI.parse(target).normalize
uri.absolute? or raise 'URI not absolute'
%w[ http https ftp ].include?(uri.scheme) or raise 'Unsupported URI scheme'
# Etchttps://github.com/ruby/ruby/blob/trunk/lib/uri/rfc2396_pars...
I haven't tested it myself, but it's worth looking at.
Original post: http://daringfireball.net/2009/11/liberal_regex_for_matching...
Updated version: http://daringfireball.net/2010/07/improved_regex_for_matchin...
Most recent announcement, which contained the Gist URL: http://daringfireball.net/linked/2014/02/08/improved-improve...
[1] http://www.fileformat.info/info/unicode/char/272a/index.htm
That is a valid domain name and should be treated as such.
DISALLOWED: Those that should clearly not be included in IDNs. Code points with this property value are not permitted in IDNs.
That said, there is a regex in RFC3986, but that's for parsing a URI, not validating it.
I converted 3986's ABNF to regex here: https://gist.github.com/mnot/138549
However, some of the test cases in the original post (the list of URLs there aren't available separately any more :( ) are IRIs, not URIs, so they fail; they need to be converted to URIs first.
In the sense of the WHATWG's specs, what he's looking for are URLs, so this could be useful: http://url.spec.whatwg.org
However, I don't know of a regex that implements that, and there isn't any ABNF to convert from there.
And why must the root period behind the domain be omitted from URLs? Not only does it work in a browser (and people end sentences with periods), the domain should actually end in a period all the time but it's usually omitted for ease of use. Only some DNS applications still require domains to end with root dots.
It's not fancy but it will essentially match any url
There are supposed URIs in that list that aren't actually URIs, there are supposed non-URIs in that list that are actually URIs, and most of the candidate regexes obviously must have come from some creative minds and not from people who should be writing software. If you just make shit up instead of referring to what the spec says, you urgently should find yourself a new profession, this kind of crap has been hurting us long enough.
(Also, I do not just mean the numeric RFC1918 IPv4 URIs, which obviously are valid URIs but have been rejected intentionally nonetheless - even though that's idiotic as well, of course, given that (a) nothing prevents anyone from putting those addresses in the DNS and (b) those are actually perfectly fine URIs that people use, and I don't see why people should not want to shorten some class of the URIs that they use.)
By the way, the grammar in the RFC is machine readable, and it's regular. So you can just write a script that transforms that grammar into a regex that is guaranteed to reflect exactly what the spec says.
Also, what exactly is the problem with email addresses? There is a very unambiguous grammar of those in the RFC, and there are lots of implementations of exactly what the spec specifies. Just because some web kiddies have made up some shit about email addresses and use that for validation, doesn't mean that postfix, qmail, or exim are written by morons.
This is exactly what Anne van Kesteren has been doing with the URL Standard: http://url.spec.whatwg.org/
> By the way, the grammar in the RFC is machine readable, and it's regular.
The RFC does not reflect reality either (which, ironically, is what you seem to be complaining about). If you’re looking for a spec-compliant solution, the spec to follow is http://url.spec.whatwg.org/.
> If you just make shit up instead of referring to what the spec says, you urgently should find yourself a new profession, this kind of crap has been hurting us long enough.
I am aware of, and am a contributor to, the URL Standard: http://url.spec.whatwg.org/ That doesn’t mean there aren’t any situations in which I need/want to blacklist some technically valid URL constructs.
WTF? What is "validation" supposed to be good for if it doesn't actually validate what it claims to? Exactly this mentality of making up your own stuff instead of implementing standards is what causes all these interoperability nightmares! If you claim to accept URLs, then accept URLs, all URLs, and reject non-URLs, all non-URLs. There is no reason to do anything else, other than lazyness maybe, and even then you are lying if you claim that you are validating URLs - you are not. If you say you accept a URL, and I paste a URL, your software is broken if it then rejects that URL as invalid.
This does not apply to intentionally selecting only a subset of URLs that are applicable in a given context, of course - if the URL is to be retrieved by an HTTP client, it's perfectly fine to reject non-HTTP URLs, of course, but any kind of "nobody is going to use that anyhow" is not a good reason. In particular, that kind of rejection most certainly is something that should not happen in the parser as that is likely to give inconsistent results as the parser usually works at the wrong level of abstraction.
> The RFC does not reflect reality either (which, ironically, is what you seem to be complaining about).
Well, or reality does not match the RFC?
> If you’re looking for a spec-compliant solution, the spec to follow is http://url.spec.whatwg.org/.
A spec for a formal language that doesn't contain a grammar? The world is getting crazier every day ...
> That doesn’t mean there aren’t any situations in which I need/want to blacklist some technically valid URL constructs.
Yeah, but blocking IPv4 literals of certain address ranges seems like a stupid idea nonetheless. Good software should accept any input that is meaningful to it and that is not a security problem. And as I said above, such rejection most certainly should not happen in the parser.
> Well, or reality does not match the RFC?
Doesn’t matter – if there’s a discrepancy between what a document says and what implementors do, that document is but a work of fiction.
> And as I said above, such rejection most certainly should not happen in the parser.
This is not a parser.
Yes and no. When there is a de-facto standard that just doesn't happen to match the published standard, yeah, sure. Otherwise, bug compatibility is a terrible idea and should be avoided as much as possible, many security problems have resulted from that.
> This is not a parser.
Well, even worse then. Manually integrating semantics from higher layers into parsing machinery (which it is, never mind the fact that you don't capture any of the syntactic elements within that parsing automaton) is both extremely error prone and gives you terrible maintainability.
edit:
For the fun of it, I just had a look at the "winning entry" (diegoperini). Unsurprisingly, it's broken. It was trivial to find cases that it will reject that you most certainly don't intend to reject. For exactly the reasons pointed out above.
"%x"
(edit: actually, feel free to do it with the quotes included if you like, that would still be matched by that regex)