HNHacker News
TopNewBestAskShowJobs

mathias

3,588 karma · joined December 26, 2009

Web standards enthusiast. https://mathiasbynens.be/
submissionscomments
mathias··on In search of the perfect URL validation regex
> > The RFC does not reflect reality either (which, ironically, is what you seem to be complaining about).

> Well, or reality does not match the RFC?

Doesn’t matter – if there’s a discrepancy between what a document says and what implementors do, that document is but a work of fiction.

> And as I said above, such rejection most certainly should not happen in the parser.

This is not a parser.

mathias··on In search of the perfect URL validation regex
But then you might end up shortening things like `about:blank` by accident.
mathias··on In search of the perfect URL validation regex
You do realize that RFC 3986 doesn’t actually match reality, right? http://url.spec.whatwg.org/#goals
mathias··on In search of the perfect URL validation regex
My exact use case was the following: the user clicks a bookmarklet that passes the current URL in the browser as a query string parameter to a URL shortener script. The validation is then performed before the URL is shortened.

In that scenario, and with the given requirements, I can’t think of a case where the validation fails. There’s no need to worry about protocol-relative URLs, etc.

(Keep in mind that this page is 4 years old — I very well may have missed something.)

> If you really want your URL shortener to reject bad URLs, then you need to actually test fetching each URL (and even then...)

I disagree. http://example.com/ might experience downtime at some point in time, but that doesn’t mean it’s suddenly an invalid URL.

> As an aside, I'd instantly fail any library that validates against a list of known TLDs. That was a bad idea when people were doing it a decade ago. It's completely impractical now.

Agreed.

mathias··on In search of the perfect URL validation regex
Yes, to reject non-URLs and also some URLs that are technically valid but that I want to explicitly disallow anyway.
mathias··on In search of the perfect URL validation regex
That was not an option in this case, as the goal is to validate URLs entered as user input and blacklist certain URL constructs even though they’re technically valid.
mathias··on In search of the perfect URL validation regex
> Deviating from the formal spec because everyone practically agrees how to do things, albeit differently than in the formal spec, is something quite different from making shit up, and actually tends to be even harder than building things to spec, as there tends to be no easy reference to look things up in, but instead you might have to look into the guts of existing implementations and talk to people who have built them to figure out what to do - and you would normally start with an implementation according to spec anyhow, and only add special cases for non-normative conventions lateron.

This is exactly what Anne van Kesteren has been doing with the URL Standard: http://url.spec.whatwg.org/

mathias··on In search of the perfect URL validation regex
The goal was to come up with a good regular expression to validate URLs in user input, and not to match any URL that browsers can handle (as per the URL Standard). I am fully aware that this is not the same as what any spec says.

> By the way, the grammar in the RFC is machine readable, and it's regular.

The RFC does not reflect reality either (which, ironically, is what you seem to be complaining about). If you’re looking for a spec-compliant solution, the spec to follow is http://url.spec.whatwg.org/.

> If you just make shit up instead of referring to what the spec says, you urgently should find yourself a new profession, this kind of crap has been hurting us long enough.

I am aware of, and am a contributor to, the URL Standard: http://url.spec.whatwg.org/ That doesn’t mean there aren’t any situations in which I need/want to blacklist some technically valid URL constructs.

mathias··on In search of the perfect URL validation regex
You forgot the most relevant spec, the URL Standard: http://url.spec.whatwg.org/
mathias··on In search of the perfect URL validation regex
As for why the trailing dot is disallowed, see <http://saynt2day.blogspot.com/2013/03/danger-of-trailing-dot....

The goal was to come up with a good regular expression to validate URLs as user input, and not to match any URL that browsers can handle (as per the URL Standard).

a.b--c.de is supposed to fail because `--` can only occur in Punycoded domain name labels, and those can only start with `xn--` (not `b--`).

mathias··on In search of the perfect URL validation regex
Exactly. The goal was to come up with a good regular expression to validate URLs as user input. There’s no way I’d want to allow alternate IP address notations.
mathias··on In search of the perfect URL validation regex
If you looked at the page before commenting you’d know that PHP’s built-in URL parser is one of the implementations that is being tested. You’d also see that one of the requirements was to not match scheme-relative URLs (e.g. `//foo.bar`) like the ones your commit fixes.

The regular expressions presented do not conform to the URL Standard or to any RFC, but rather to the list of requirements on that page.

I applaud your work in improving PHP’s built-in URL parser.

mathias··on Everything You Need to Know About the CSS will-change Property
The post links to it as well.
mathias··on Flickr: Invitations disclosure (resend feature)
Which guideline in particular? “Otherwise please use the original title, unless it is misleading or linkbait”? The new title is definitely not as clear IMHO.
mathias··on Flickr: Invitations disclosure (resend feature)
Looks like someone just changed the title of this HN submission. For the record, it originally said: “Full name and email for every Flickr invite ever sent can be viewed by anyone”, which was accurate at the time of posting.
mathias··on Flickr: Invitations disclosure (resend feature)
Minor correction: it was “d4d1a179c0f3” who reported and suggested that solution, not me (although I agree with their proposal). I’m just the guy who posted this on Twitter (https://twitter.com/mathias/status/452714683628527616) and Hacker News.
mathias··on Hide this in your coworkers' JavaScript code tomorrow
For more, see the evil.js project: https://github.com/kitcambridge/evil.js

Also note that `string.split('').reverse().join('')` is not a very good way to reverse a string in JavaScript. See http://mathiasbynens.be/notes/javascript-unicode#reversing-s...

mathias··on PBKDF2+HMAC hash collisions explained
Some PBKDF2-HMAC-SHA1 collisions that start with the URL of this Hacker News thread:

    $ ./brute-force.py 'https://news.ycombinator.com/item?id=7465849#'
    https://news.ycombinator.com/item?id=7465849#aaaaaaaaaaaaabufhkcn 💥 VID9Qf*}8w9sBtI@TK5k
    https://news.ycombinator.com/item?id=7465849#aaaaaaaaaaaaabwwhwta 💥 Bgvi4F~6M#utux ~\H4m
    https://news.ycombinator.com/item?id=7465849#aaaaaaaaaaaaacjypwdz 💥 Tn/ZyN'Zs6?d){Ytd?}i
    https://news.ycombinator.com/item?id=7465849#aaaaaaaaaaaaaewyefbo 💥 r;9U\DZh!2de#bt <2*o
    https://news.ycombinator.com/item?id=7465849#aaaaaaaaaaaaaezrsysb 💥 wl)@eb.1Avvp?|NtT(\c
    …
mathias··on Speaking JavaScript
If you’d take the time to read the free online copy (which OP links to) you’d see that there are plenty of “pure examples” in the book. In fact, there’s a code example for every two paragraphs of text, or so.
mathias··on Speaking JavaScript
“I didn't remember the exact wording, but he/she mentioned something in the lines of Javascript is better written than spoken.”

In that case, you’re gonna love this book. There’s a small code example for every two paragraphs or so, which really helps make things clear. (Disclaimer: I was a technical reviewer for this book.)

mathias··on Speaking JavaScript
What did the original post say?
mathias··on To close or not to close – Void HTML elements
I used the SGML NET trick a few years back in an attempt to create the shortest possible valid HTML documents for different versions of HTML: http://mathiasbynens.be/notes/minimal-html

Note: “valid” here is defined as “theoretically valid as per the relevant spec” and doesn’t reflect what browsers actually support(ed).

mathias··on To close or not to close – Void HTML elements
`<script src="foo" />` only works the way you’d expect it to in XHTML. Proper XHTML, that is — served with the correct `Content-Type` header. http://mathiasbynens.be/notes/xhtml5
mathias··on To close or not to close – Void HTML elements
From the article:

“It is not, and has never been, valid HTML to write `<br></br>`.”

Sure, but note that it is perfectly valid XHTML (which is a form of HTML).

Oh, and `<script src="foo" />` actually works the way you’d expect it to in XHTML.

Don’t use XHTML though.

mathias··on Apple Removes Shadow DOM from Safari
The 8.8 million lines joke refers to <http://techcrunch.com/2013/05/16/google-has-already-removed-....
mathias··on Opera is really nice.
What is your definition of ‘non-free browser’?
mathias··on Opera is really nice.
Are you comparing a mobile browser with a desktop browser?

Note that Opera’s “Off-Road mode” in present in Opera for Desktop as well; it’s not a mobile-only feature. Here’s a screenshot that shows how to enable it: http://i.imgur.com/rvGkQrI.png

mathias··on Opera is really nice.
Note that Opera’s “Off-Road mode” in present in Opera for Desktop as well, so it’s not a mobile-only feature. Here’s a screenshot that shows how to enable it: http://i.imgur.com/rvGkQrI.png
mathias··on Opera is really nice.
> In the mobile world they offer access to their compressing proxy built-in

Off-Road mode in present in Opera for Desktop as well. Here’s a screenshot that shows how to enable it: http://i.imgur.com/rvGkQrI.png

mathias··on Opera is really nice.
Some cool features I haven’t seen in Chromium yet:

Lazy session loading: when restoring browser state from session, only the active tab will be loaded. Remaining tabs will be loaded when activated. Enable here: opera://flags/#lazy-session-loading

Extended lazy session loading: changes mode of operation of lazy session loading so that all tabs are gradually loaded in the background. Enable here: opera://flags/#extended-lazy-session-loading

I use my browser tabs as a to-do/to-read list – bad habit, I know. At any given time I have at least 40 tabs open. For people like me, features such as the above boost browser startup time significantly and help consume less memory.

← PreviousPage 2 of 7Next →