By "diversity," do you mean something like "entropy?" Like maybe
H_s(x) := -\sum_{x \in X_s} p(x) log(p(x))
where X_s := all s-grams from the training set? That seems like it would eventually become hard to impossible to actually compute. Even if you could what would it tell you?Or, wait... are you referring to running such an analysis on the output of the model? Yeah, that might prove interesting....