This is great.
I wish we could do our own arbitrary style analysis on the data set sort of the way one can do a factor analysis on a portfolio.
I would look at words that are common between Marukami and McCarthy compared to the rest of the corpus for instance.