Huge topic, but briefly, we use a combination of RSS, scrapers, and official APIs to obtain article/post text. As far as clustering, it's a combination of text analysis, link analysis, and editorial input. I don't expect us to open source much of our code unless we find a new source of revenue!