One common way to test it is just to pass ipv6 url: http://[f021:d981:b487:e57d:193e:550e::]/
One common way to test it is just to pass ipv6 url: http://[f021:d981:b487:e57d:193e:550e::]/
RFC 3986 Appendix B [1] "Parsing a URI Reference with a Regular Expression":
The following line is the regular expression for breaking-down a well-formed URI reference into its components.
^(([^:/?#]+):)?(//([^/?#]*))?([^?#]*)(\?([^#]*))?(#(.*))?
scheme = $2
authority = $4
path = $5
query = $7
fragment = $9
Let's test your URI with this regex, shall we? [2]
$2 (scheme) = http
$4 (authority) = [f021:d981:b487:e57d:193e:550e::]
$5 (path) = /
Seems correct to me.[1] https://datatracker.ietf.org/doc/html/rfc3986#appendix-B
https://github.com/bkaradzic/bx/blob/0b001f5f36579e8aea07efa...
That's a problem orthogonal to URI parsing.
You parse the URI with the RFC 3986 regex, which gives you the components: scheme, authority, path, query, fragment.
You're then free to parse any of the components according to your own bespoke rules, e.g. the query string often follows the key=value&key=value&... pattern.
A parser that assumes the input to be already valid is usually not enough for most applications
according to that regex this is a valid url
__..__..%%%zz..
Correctly. It is a valid relative URI, whose path is "__..__..%%%zz..".
also this one is valid too for the regex
%_%://%_%.%_%/%_%?%_%#%_% ht tp://[\x00::gg]\t/foo\nbar/%/%QZ/%F?q=<\"|{}>##Does it need to be human readable, does it need to work across all platforms. Does the it need to be secure. These things change everything. Speed, reliability, security pick one.
Your point about being compliant with the real spec is the difference between a 20 line scannf and and a 1000 line function. Yeah. Ha.