As a side note, the new HTML is way more complicated and much harder to parse than before - I know the aim isn't to help parsing for content, but I was still saddened to see how it's ended up (a bit of a mess imo - hard to distinguish actual article content from other things).
If anyone knows a reliable and public way to access the content before the "web rendering" layer, that'd be very handy!