I extracted the inner text inside a <p class="foo"></p> using Python & BeautifulSoup, print to stdout, then cut, sort, uniq, and sort again.
It looks like the counts from this voodoo are incorrect: they are all twice as much. However, proportions are still correct.