Basically the first step would be shingling the text (choosing a sampling domain) and generating a MinHash struct (computationally cheap) which can then be used to find the "similarity" between sets, or, the "Jaccard Index."
If you're clever about this, you can use HyperLogLogs to encode these MinHash structs gaining a great deal of speed with a marginal error rate, all while allowing for arbitrary N-levels of intersection.
If you're looking to build a model to analyze two (or N) text bodies for stylometric similarities, I'd approach the problem in two steps:
1) Minimize the relevant input text.
- Use a bernoulli/categorical distribution to weight words according to uniqueness--NLP and sentiment extraction techniques may also help
- Design a markov process to represent more complex phrasing patterns for the text as a whole
- Filter by a variable threshold to minimize the resulting set of shingles/bins/"interesting nodes" into a computationally-manageable #
2) Use an efficient MinHash intersection to compute a similarity vector (0-1) for the two texts.
I think given the prevalence of training data (I mean, what's more ubiquitous than the written word...) you could probably tune this to a reasonable accuracy and efficient complexity.
Just a 5m thought exercise, but if anyone else has ideas I'd be curious as well :)