Show HN: Tuchu – Automatically highlight the important parts of a document
tuchu.app
tuchu.app
I created Tuchu because I wanted to increase my reading efficiency. It is a tool that automatically highlights the important parts of any document. Most documents take about two to three seconds to process. It's directed at students, researchers, or anyone with a reading list and little time for that matter.
During my studies I had to go through a lot of literature, for example when I had to select relevant material for my thesis, or when I had to familiarize myself with a course's reading list. Tuchu helped me to get up to speed in these cases. What started off as a command-line Python script is now a web application that does its analysis without any back-end. I don't get to see your documents.
The underlying algorithm that selects what's relevant is called TextRank, an unsupervised summarization method [1]. It models a document (or a collection thereof) as a fully connected graph. Its nodes are parts of the text — I use sentences — and the edges between them are weighted by a similarity measure, in my case simple word overlap. The subset of sentences with the highest PageRank are then highlighted. For good measure, I also highlight sentences that contain signal words that — in my academic experience — signify importance.
It's important to note that Tuchu is not a substitute for doing your own reading. It could make you a faster reader by directing your attention to the important parts, but you'll still have to ponder about the true essence of a document yourself.
I uploaded a PDF version of a Wikipedia article to see what was selected, and at a quick glance, it's not obvious to me that the most important parts of the article have been highlighted. On the other hand, it's not obvious that trivial parts have been selected either -- leaving me intrigued to look further.
Nowadays a lot of bloggers/reporters keep running round in circles before coming to a point. I actually wrote a post related about this a while back [0]. With browser extension this can save lot of time of readers. Heck I would pay for such a service!
At this point you could print the page to a PDF and then upload it. A browser extension for webpages (as suggested by others) is a good idea! I'll look into it.
Cool! Would love to hear more. Perhaps you could reach out at hello@tuchu.app? :)
This is weird, as Tuchu means highlight in Chinese.
Thanks for your feedback. In theory all languages should be supported since TextRank should be language-agnostic. Having said that, I think that my current sentence similarity strategy (by word overlap) is not a good fit for "symbol rich" languages such as Mandarin.
I am not Chinese (or from Asian descent) myself. Tuchu comes from some rainy Sunday Google translating. ;)
I like it. it's got potential. With north of 5000 pdfs in my library I'm extremely open to tools like this, and the methodology seems pretty sensible to me. But at present it feels kind of random - good at picking out summary sentences of what a document section will cover, or emphatically stated conclusions, bad at highlighting necessary context. I wonder if it might be better trained on clauses rather than full sentences, although that's probably significantly more work.
I'm sorry I don't have a more positive review, but I think it's a bold attempt even if it falls short, and want to see where it goes. Even as is, I can see myself using it sometimes as its selections offer an interesting alternative to my own skimming/emphasis preferences. Very impressed with the clean no-BS user interface and fast performance.
In my opening post I said that doing your own reading is still a requirement, but it might be good to also mention this on the site — at least until I find a way to improve the algorithm. One way could be leveraging pre-trained word embeddings, but this would require a server or downloading a large blob to the user's device beforehand. In any case Tuchu wouldn't be as fast that way.
https://ec.europa.eu/transparency/regdoc/rep/1/2020/EN/COM-2...
Good idea! I'll note it down. :)
In any case, the paper I submitted is one I coauthored, so I like to think I'm a reasonably good judge of what's important. Maybe the tool just isn't a good fit for my field or my writing, but the highlights appear to be essentially random.
Thanks for reporting this! All feedback is really welcome. Firefox is supported. The error you encountered originates from the underlying PDF renderer which, depending on the browser you use, sometimes throws an uninterpretable error. It's on my list...
With respect to your highlighting results: I am aware that it can be a hit or miss at this stage. I've had really mixed feedback so far and I do have a theory that the quality of results may depend on the kind of writing (which is odd, bit I digress). If you could send your document to hello@tuchu.app that would be really useful!
One cool thing is that it highlighted both the same sentences in English and French. I presume you translate the text before analyzing.
[1] https://www.math.mcgill.ca/darmon/theses/leahy/thesis.pdf
There's no translation being done at all! It's completely language-agnostic (apart from some hardcoded signal words, which I only noted down in English).
Really large documents such as entire books or theses are supported, but as you've experienced it results in a far from ideal experience at this point. Thanks for reporting this and linking to the document, now I know that this is a problem worth addressing. :)
EDIT: I removed my personal email from this message.
I did it by hand. The hardest part was finding a way to create the sentence similarity matrix in a fast way. I solved that problem by creating a Matrix class around Javascript's typed arrays.
For the complete picture, the app itself is written in Angular because I'm familiar with it (probably overkill, really) and I'm using Mozilla's pdf.js [1] to render documents.