High-reproducibility and high-accuracy method for automated topic classification
amaral-lab.org
amaral-lab.org
I'm working on a project to make (EU-) law more accessible. So if anybody here knows good methods to visualise/summarise long legal texts (30-300 pages) you could do something for humanity by posting a reply.
(Word clouds just don't cut it in these cases.)
1. Split your text into sentences
2. Remove stopwords (keeping a copy of the original sentence)
3. For the remaining words, calculate the synthetic TF-IDF score of all the words in the sentence (tf-idf each word then sum them).
4. Keep the n-highest scoring sentences (in the order they appear), or all the sentences with synthetic TF-IDF scores above some threshold.
There's your summary.
BTW, I love the highlight-in-the-original-text option.
> How did it work? Any interesting tweaks you can share?
For being such a naïve approach I think it works fantastically. It works better on some types of texts (newspaper articles, research papers) than others (poems, fiction). It was my first real programming project so the code has a certain sentimental meaning for me :)The analysis code (https://github.com/peterldowns/bookshrink/blob/master/analys...) is very short and the comments include my thoughts on certain tweaks / approaches. Some quirks of the implementation: uses regex for sentence splitting (really!), doesn't perform any stemming, and weights proper nouns heavier than other words.
> BTW, I love the highlight-in-the-original-text option.
Thanks! The eventual goal was to build a tool that would make essay grading easier for teachers, although I never got around to it.There's a bit of art to it for sure that can improve the results. You may also have to do some pronoun substitution in the summarized sentences (and then decide to do that before or after calculating the synthetic score) so they make more sense.
I find it's hard to tell how well all this will work until you just do it. Not every kind of text works equally well and the only proof that it's really working is "does the summary make sense or not?"
It's also possible that for long structured documents like laws or contracts, that you don't want to summarize the whole thing, but treat major sections like different documents and do intra-document summarization to maintain understandability.
Here's one that does something kinda like what I was writing about above.
note I think the correct measure is not TF-IDF, but TF-ISF (Inverse sentence frequency) for single document summarization, but I might be wrong.
related is the idea of "stemming" which uses an algorithm to try to reduce inflection, to find a common form of a word that various versions come from. Porters algorithm is a well known stemming algorithm. However, sometimes you end up with weird "non-inflected" tokens at the end. (e.g. 'enhancement' might become 'enhanc')
However, lemmatization is considered "better" in that it uses a dictionary of inflected forms that map back to the non-inflected form. So in theory, if the dictionary is comprehensive, you can properly replace inflected forms with their correct non-inflected forms. (e.g. 'enhancement' -> 'enhance')
If your dictionary isn't comprehensive and comes across a token it doesn't recognize, you can try falling back onto a stemming algorithm.
http://www.cs.ubc.ca/labs/imager/tr/2014/Overview/
It includes a number of the data mining techniques others have mentioned and implements some nice design guidelines from visual analytics.
And it's not like the people in this field haven't been aware of network-oriented methods. But rather than using community-detection as a mechanism for topic discovery, instead people either focused on networks among topics to see how topics are related, networks among authors such that social network information informed topic discovery, or networks among documents where link/reference information was explicitly part of the model.
These authors seem to get solid results in part by having totally different values/aesthetics. Unlike the Bayesian nonparametrics people, they clearly don't care about picking arbitrary, inflexible parameters (e.g. the 5% threshold), nor do they want their model to have a clear, generative form, nor are they particularly concerned about having a new algorithmic insight (since they throw their hard work to InfoMap, and discuss none of its details), nor do they attempt to advance the expressiveness of their topic model (they proceed with the most basic bag-of-words model available). But it does seem like they get good results on the basic task with a very pragmatic, pipeline approach.
I'm not entirely familiar with LDA, but from what I was able to understand from their intro, it feels like their LDA application could have used some feature selection.
The journal they published in, Physical Review X, is a newer open-access journal from APS (along the same lines as PLOS ONE or Nature Scientific Reports). I think it's great but not everyone agrees. To read more on the debate around the open-access phenomenon look at http://blogs.berkeley.edu/2013/10/04/open-access-is-not-the-... and http://www.sciencemag.org/content/342/6154/60.full
In my experience, finding out when a paper was published is often not too difficult. In most cases, Google makes it easy to find the Journal issue and/or the conference proceedings of the paper. Or you find some third paper that contains a reference with date information that you can then use to double check.
What about that sentiment analysis NLP tool that someone posted on HN last year? That was also very good.
Deleted comment
Who needs to understand them? The algorithm? What does that even mean? And if you mean the authors need to understand how the algorithm works: Wrong again. They probably do, but even if they wouldn't their algorithm might still classify correctly.
Science works by separating out the disciplines. Frankly I think "defense against the terminator scenario" could and should be a scientific field on its own at this point on the level of solutions to global warming.
This is interpreted as a side effect for now. That is, until you tell the computer to do something based on the topics.
If the output of AGI demonstrates intelligence in finding a solution, whether it had "understanding" or not doesn't matter. The only thing that matters is its power to turn inputs into outputs.
The Turing test crystallizes this WRT human language interaction. Assessment of the intelligence of a machine doesn't depend on understanding, consciousness or any of the other baggage dragged in.