XHTML1 depends on DTD for entities, which means that parser must fetch DTD from w3.org (from heavily-hammered server) or support DTD catalog and have one set up in the system (rarely available out of the box, not always possible to do). IMHO that breaks "works everywhere out-of-the-box" promise, and sometimes it makes it easier to just use HTML parser.
It's been 11 years since XHTML recommendation was published, and you still can't rely on XML parser for reading content on the web. Even documents labelled as "XHTML" are often ill-formed (and almost all of them are sent as text/html). Even if we were moving in the right direction, we were moving too slowly. Now that HTML5 parser is specced, it may be quicker to add it to popular languages than turning whole web around.