In search of the perfect URL validation regex
mathiasbynens.be
mathiasbynens.be
http://www.foo.bar./
http://10.1.1.1
http://10.1.1.254
http://10.1.1.0
http://10.1.1.255
The top one is just a fully qualified domain name, and the rest could be valid host addresses depending on your subnet size.This shouldn't fail. It resolves to an IP owned by The Coca Cola Corp (not an internal IP)
It should also start with a 0 to be "correct". Which in this case "correct" means the libc inet_aton() function will accept it and convert it to an ip address. The octal portion is nonstandard technically for uri's and is more an artifact of the libc the software you're using is sitting on.
Regardless, there are valid dotless decimal, octal, and hex representations for every IP address. Library support is variable.
It's valid, but the only people who ever use dword IP addresses are pentesters and their less scrupulous analogues. I have no issue disapppointing them.
There are valid dotless (and dotted) octal and hex representations too. They should also fail, I think.
I also don’t want to allow every possible technically
valid URL — quite the opposite.
And dword and octal IP URLs are definitely in the realm of technically valid but practically not.Well, I trudged through the RFCs in ~2010, so it's possible that it's been updated since then.
You wouldn't write a normal function in one line with no comments, why do it with regex?
URIs are one of those things that make you go, "Oh! That should be easy!", and then a week later you're walking around looking for puppies to kick.
For example, http://1249711460 resolves just fine in Chrome.
Certainly a witch nest. :]
74.125.21.100 => 100x1 + 21x256 + 125x256^2 + 74x256^3 => 1249711460
In all seriousness, this is why I don't like the IETF's documents. They write in a verbose way and then don't even provide a reference implementation, in this case, a reference regex that would have done away with lots of dispute and ambiguity. It is my opinion that, in practice, your specification is inherently broken if you can't provide a reference implementation.
(No, BNF doesn't count as reference "implementation". Who uses BNF in their programs to validate strings anyway?)
And that is why IETF RFC are so verbose. Because there is politics in them. That is why you can't provide regexps.
url because of politics are not context free. https://www.cs.rochester.edu/~nelson/courses/csc_173/grammar...
So a URL because of politics is a context full grammar.
And regexp cannot parse anything else than context free stuff.
To be honest I doubt anything useful is context free.
I doubt, therefore, that anything useful can be parsed with regexp... Except float, integer, and other basic types that are useful to build a context full grammar... But for this RFC should separate standards in context free stuffs (rules/context free stuff) and config file for the political/commercial stuffs (context that changes meanings of atoms being parsed for illogical reasons).
The problem is politics is fucking hard to normalize, we have no BSML yet.
[1] https://en.wikipedia.org/wiki/Regular_grammar
[2] https://en.wikipedia.org/wiki/Regular_expression#Patterns_fo...
I've written a lot of web facing software that accepts URLs from the untrusted masses and ultimately makes requests to them if they are "valid." The lesson I've learned is simple. Regex's are terrible for this task because there are a ton of things you check and lots of normalization you need to do. Instead, do this as a function
I've evolved mine over the years, and my use case is semi-specific: Given a string, validate it as a fully qualified HTTP/HTTPS URL that doesn't have credentials and isn't trying to point my software toward the internal network or localhost. It looks like this:
- Use system/framework library to create Uri object from source string. All your checks will be consulting this object's properties, not looking at the source string
- Is scheme HTTP/HTTPS? If no, stop
- Did they supply user:pass@ in URL? If so, stop, and yell at them for putting usernames and passwords into a random site on the Internet.
- If hostname is an IP address, normalize it to dotted decimal quad IPv4 or IPv6 (no octal obfuscation for you!), and test against private IP space ranges or loopback. If private or loopback, stop
- If hostname is an actual hostname, normalize it + de-puny code it, and check for localhost aliases. If local, stop (you can also do a DNS lookup and make sure you someone isn't trying to return private/local IPs to bypass your checks)
At this point, you have a syntactically valid, fully qualified URL pointing to a public facing web property accessed via HTTP or HTTPS.
You don't have to worry about TLDS or the like. At this point, you can do additional DNS checks, check the domain against lists of bad actors, whatever else you want to do. You can try and be smart and do things like, "if supplied URL wasn't fully qualified, prepend http:// and try validation again" to avoid user error. Pretty flexible.
This is more rigorous than a simple regex and way way way easier for another developer to read and understand what is going on.
FWIW, browsers don't send the user:pass in the URL - they automatically marshal it into the Authorization: Basic header. Obviously sending these over HTTP is still dumb, but they're not (typically) logged in plain-text in browser history/server request logs.
At least once a month I get someone giving me a URL with embedded credentials to a dev or staging environment for an Alexa top 50K site. Things like https://qa:tester@dev.major-ecomm-site.com/blah. It's pretty terrible, but makes for an interesting sales call :-)
http://foo:bar.com\@example.com
Chrome/Webkit browsers will go to `bar.com` but Firefox will end up at `example.com` for the same URL! There's a fantastic book called "The Tangled Web" that has lots of examples of these pitfalls.
Is this because people think of a regex as an atomic blackbox instead of a program or function that can be read and modified? For example when a regex is incorrect, they don't say "This regex is incorrect. How can it be corrected?" like they would for a program or function, but "This is the incorrect regex. Which one is the correct regex."
Free-spacing mode (where you can add whitespace, newlines and comments to a regex) would at least help a little. (Although even then I think a function with early returns is often more appropriate. Possibly using very small regular expressions for some of the individual checks.) But I rarely even see free-spacing modes mentioned. Maybe because most programmers using regex actually prefer an atomic blackbox?
Validating URLs is also something you shouldn't have to implement yourself.
> This requirement is a willful violation of RFC 5322, which defines a syntax for e-mail addresses that is simultaneously too strict (before the "@" character), too vague (after the "@" character), and too lax (allowing comments, whitespace characters, and quoted strings in manners unfamiliar to most users) to be of practical use here.
However as others already pointed out, it's better not to use a regex for it, but a proper library for your language which would bail out as soon as it hits something invalid. With the hope that the library for the language does support that already.
The way it defines passwords is wrong, so you can trick it into accepting almost anything by putting the domain somewhere else:
re_weburl.test("http://127.0.0.1/")
=> false
re_weburl.test("http://127.0.0.1/@example.com")
=> true
re_weburl.test("http://999.999.999.999.999/@example.com")
=> true
oops. I disagree that it should be rejecting rfc1918 addresses anyway, because this makes it less useful in an intranet context, where you want those to work.There's also an apples-to-oranges comparison going on here. The Gruber pattern is not for validation, but for detecting url-like-things in text, which is why it excludes a whole bunch of punctuation chars from appearing at the end - when I say 'google.com.' in text, I mean 'google.com'.
Edited to add: I misremembered suggesting trailing punctuation exclusion to Gruber, what we discussed was xxx.xxx/xxx as an alternate pattern, catching protocol-less shortened urls in tweets.
This would be much easier to do with a context free grammar or any decent parsing library.
https://github.com/nisavid/spruce-iri/blob/master/spruce/iri...
Yes, I wondered about this myself, http://dk. redirects to www.dk-hostmaster.dk.
It would be interesting to receive mail from root@com or whatever :)
$ dig va MX
...
;; ANSWER SECTION:
va. 3599 IN MX 10 mx12.vatican.va.
va. 3599 IN MX 10 mx11.vatican.va.
va. 3599 IN MX 100 raphaelmx3.posta.va.