If the data is available as a bulk download, you're all set to move to the next step. If not, next best is if they offer an API. In that case, learn how to use Python or something to pull each document using that API, and storing it somewhere (either the filesystem, or in a database). If not, get the book 'Web Scraping with Python' and use that.
Once you have the stuff together, Udacity has a gentle introduction on cleaning document-ish data (using JSON/Python and Mongodb): https://www.udacity.com/course/data-wrangling-with-mongodb--...
Then the analysis starts. If it were me, I might start by splitting documents into chunks (paragraphs?), and classifying them somehow. Maybe use NLTK: http://www.nltk.org/book/ch06.html
Let us know how you get on :)