A new method to validate URLs in JavaScript
stefanjudis.com
stefanjudis.com
Is it? According to https://www.w3.org/TR/2011/WD-html5-20110525/urls.html, a URL is a valid URL if at least one of the following conditions holds:
• The URL is a valid URI reference [RFC3986].
• The URL is a valid IRI reference and it has no query component. [RFC3987]
• The URL is a valid IRI reference and its query component contains no unescaped non-ASCII characters. [RFC3987]
• The URL is a valid IRI reference and the character encoding of the URL's Document is UTF-8 or UTF-16. [RFC3987]
Here's how I do it: https://gist.github.com/eligrey/443d51fab55864005ffb3873204b...
As you can see, this implementation fails for URLs with - as their host. There should be a cleaner solution here that doesn't involve creating new URL instances with test scaffolding (the input base URL).
https://developer.mozilla.org/en-US/docs/Web/API/URL/canPars...
I disagree.
Your approach does not address[0]:
By convention, domain names can be stored with
arbitrary case, but domain name comparisons for
all present domain functions are done in a
case-insensitive manner, assuming an ASCII
character set, and a high order zero bit.
Nor URL's with user and/or password elements[1].Nor the different, yet equivalent, path encodings such as "/has+space" and "/has%20space".
Nor does it normalize path segments before determining value equality, such as "/foo/bar" being equivalent to "/foo/blah/../bar" as well as be equivalent to "/../../../foo/../foo/blah/../.././foo/./bar".
So no, validating absolute URL's is not "easy."
0 - https://datatracker.ietf.org/doc/html/rfc1034#section-3.1
1 - https://datatracker.ietf.org/doc/html/rfc1738#section-3.1
My linked utility is specifically for validating that a URL is valid (won't throw an error when passed to the URL constructor) and doesn't need additional encoding. This helps with my use case which is a 'create URL classification' UI that allows users to input URL matchers in a multitude of formats.
For additional context, some inputs are invalid even with a base URL. e.g. new URL('//:0', 'https://-') will throw an error.
Your first critique doesn't seem relevant as this is mostly for checking if a URL is 'valid' (i.e. doesn't throw an error when used).
Also, for your second critique, the username + password is actually part of the origin as used by both of my snippets. For example, isValidURL('https://a:b@c.d/') and isValidURL('https://a:b@[::1]/') both return true for me.
Do you have a URL that gave a bad result? If so, feel free to mention it here and in the comments for the gist so that users of my snippet can be made aware of its limitation.
So maybe the post should have called it what it is: validating parse-ability, since the method itself is called "canParse()".
Url validity isn't dictated by ICANN.