the algo is still kinda "dumb". It basically tries to get top N keywords - most frequent non stopwords, stems them and goes through the sentences looking for which of them (upto max summary length) contain the keywords. I'll keep whittling at it over nights/weekends to see if I can make it more "semantically aware"
edit: some work is needed on the tokenization also. currently, I don't preserve non-period punctuation.