I'll withstand my statement: model based on a corpus of PR, scholar, licenses and the like texts.
If they are into real statistical NLP.
Or just esthetic rules + word dictionary.
Or just esthetic rules + word dictionary.
But I'm starting to think a rule-based lexicon isn't out of the question, given these >1 scores on some texts.