* The text mentions OCR, but the screenshots show documents with fidelity far beyond what I would expect via scanning. I would guess that the screenshots in actuality show PDFs that include character and layout information, i.e. they don't simply contain scanned images. If my guess is correct, why is OCR needed?
* How does segmentation contrast with layout analysis, or are they synonymous?
* I know a lot of work has been done on layout analysis in commercial off-the-shelf OCR software. How do these results (up to but obviously not including the summarization itself) compare? Or, how would you expect them to compare?
Thanks!