A Gentle Introduction to Text Summarization in Machine Learning
blog.floydhub.com
blog.floydhub.com
For training, I wonder if NER can be abused to produce a good sentence segmentation model. My idea is label the first word of every sentence as NEW_SENT and every other word as nothing. Then the model would learn which words start a sentence in a document. I haven't tried it nor know if anyone else has, but I keep meaning to try.
In those cases, the quality will depend heavily on the performance of the tokenizer as well. For Chinese you don't need a stemmer, since it's an analytic language, but Japanese is agglutinative and German is synthetic, so stemming is required for those.
A true pictorial language would convey most meaning through the symbols themselves, and I don't think any modern languages fulfill that definition. Maybe sign languages are the closest thing we have to pictorial language, in terms of the way some things are expressed symbolically?
(don't know much about NLP beyond surface level software experience but have an amateur interest in Japanese)
Ya, that's what I was trying to get at when I mentioned summarizing paragraphs into sentences. I think we're on the same page there. :)
The results can be surprisingly good, even for such a basic algorithm.
For me this technology is still in the "unusable" phase, and urgently needs more work.