As DOM query:
document.getElementsByClassName("title")[0].parentElement.getElementsByTagName("a")[1].href
This will break:* When title element no longer has "title" class.
* When title is no longer a sibling of link.
* When link is no longer 2nd link of its parent.
As regular expression:
document.documentElement.innerHTML.match('td class="title">.*a href="([^"]*)"')[1]
This will break:* On any white space change.
* On any new attributes on td or a.
* When ' is used instead of "
* When href includes escaped "
* In most cases when DOM query will break.
Many of those can happen without any server-side changes. It will sometimes works sometimes won't - making it hard to test.
There are cases when regular expression will break less often than DOM but DOM is easier to reason about, more predictable and has less corner cases.
1. Those hosts are still broken. The next person will have to jump the same hoops to support them.
2. You parser is very permissive. It will encourage people to create even more broken implementations.
3. The specification of this protocol is now worthless. There is no way to safely add new functionality. Any new element or attribute can break those regexes. Everyone has to take every implementation into account.
4. You are probably missing some corner cases like CDATA elements or quoted characters.
From your point of view, it probably makes sense to support even broken sites. But, you are helping to create next HTML - where every implementation works differently and you have to test everything on every browser.