This seems like a convoluted method to achieve the end goal, straight up regex parsing pages seems like a last ditch effort not the up front start.
OTOH this is just hearsay and I never really looked into it. So maybe this was the last ditch and OP just did not mention the failed attempts before?
All the important details are on the wiki page: https://www.mediawiki.org/wiki/Parsoid
Disclosure: I work for the Wikimedia foundation but not directly on the parsing code. Anything I post on hacker news does not represent the views of my employer.