SumatraPDF Reader
github.com
github.com
Since Adobe is pushing a more aggressive stance for monetization of Acrobat, I am trying to replace selected PDF workflows with OSS. Here are some of the tools I use.
qpdf
removing passwords, unlocking PDFs, conversion
install in WSL with apt-get install qpdf
remove password with qpdf --decrypt --password="" input.pdf output.pdf
PDF4QT - Open Source PDF Editing
Deleting, Sorting, Extracting Pages
Currently, no choco release available, must be installed manually from PDF4QT/releases
Inkscape, LibreOffice Draw
editing PDFs, adding text
Mupdf
Command line tool and Python package for parsing, filling forms, adding text
SumatraPDF
Viewing of PDFs
pdfplumber
Awesome python package to extract tables from PDFs into data pipelines. Use with Jupyter LabI got all excited - then realised "signing" just means inserting a picture. Notably absent are open source tools for digitally signing and verifying PDF's. Apparently pdftk does it in a paid version.
It's funny in a way - in this thread we have people wanting ways to modify a PDF. Yet to me, being any to prove it's not modified (eg, it's statement provably issued by some bank saying they transferred funds to my bank on behalf of person XYZ) is far more important. Instead we have companies offering paid "document signing services" which are built on sand - you can easily forge / modify any signed document they issue.
> winget search qpdf
Name Id Version Source
------------------------------
QPDF QPDF.QPDF 11.6.3 wingetThere are many wonderfully weird PDFs and epubs out there, but we do our best to fix issues. :)
[1] https://github.com/sumatrapdfreader/sumatrapdf/issues?q=is%3...
[2] https://github.com/sumatrapdfreader/sumatrapdf/issues?q=is%3...
[3] https://github.com/sumatrapdfreader/sumatrapdf/issues/2752
[4] https://github.com/sumatrapdfreader/sumatrapdf/issues/3761
I've also used poppler's pdfimages but I'd prefer like something less buggy for my use case; any version I've tried had problem with one pdf made by Adobe InDesign.
Also, tesseract allows creation of a pdf from the images with the embedded OCR text. It is also built in in the k2pdfopt.
It's like PDF4QT but works and feel better to me.
good luck finding it
> good luck finding it
It's still available on the trackers.
pdf-view-themed-minor-mode
It matches the pdf style/colors to your emacs theme! Sort of like a dark reader for pdfs, but it automatically adjusts to any theme based on some good but likely imperfect heuristics.
> And yet I do know that you can write complex, relatively bug free code without tests, because I did it.
> I do know that you can write complex, relatively bug free code without anyone looking over your code, because I did it.
> If no one uses your app then who cares if it crashes.
> If many people use your app and it crashes, they’ll tell you and then you’ll fix it.
Those four statements are contradictory. What they're saying is not that you don't need testing or code reviews, but that you can get your users to test for you.
I figure the author probably does test their code (everybody tests, even if that just means running the app), but not rigorously or in a way that you could say gives one the security of regression tests.
No-one worth discussing the issue with claims that it's impossible to write complex code without automated testing. I'm a huge proponent of automated testing, and I wrote a relatively large, cross-platform renderer without a single automated test back in the late 90s/early 00s ... it just took a long time, and I became increasingly terrified of making changes.
Edited for formatting.
What I was trying to say is: there's dogma about tests and code reviews.
At Google you would get fired for suggesting skipping code review.
Even at smaller Silicon Valley companies (smaller == less than 10 devs) it's unthinkable to not do code reviews. I haven't worked outside SV so it might be different.
That's the dogma.
My point is that maybe we should apply a bit of common sense on top of that.
I'm not saying Google should stop doing code reviews - the cost (to Google) of google search breaking is so high that you do 100x more than just code reviews.
But maybe those smaller companies don't need to dogmatically review the checkin for a documentation fix.
There's a good rule of testing top level behaviors described in this talk [1]
For code reviews, it's about knowledge handoff. No one disputes you can write great code alone. The problem is that singular geniuses writing functional but unmaintainable code only they understand and then getting hit by a bus or changing jobs is a real issue.
It's a mess. We've been paying them serious money for a product. We've never been warned that their product isn't finished yet or that we're the beta testers for the product they'll sell to other clients. Or that we have to invest our own personal and their time to fix their problems and talk to their useless support.
This has become a pattern and I'm done with it. We are slowly moving back to older and larger companies who actually do their work properly before they roll out products and updates.
I know CGM from medico...their KIS is a nightmare. We have to communicate with it in hospitals. What an ugly monster and somehow no hospital IT is able to admin it properly.
Which of their products are you considering?
What Data Warehouse software do you write for RIS? Maybe I can use it :D
...and yeah Radiology communities are rare. I'm still looking for one...if I have time.
The DWH software writes snapshots of db_direct into a temporal DB (implemented in Postgres using multiranges) and then uses dbt to transform the data into usable tables. Right now, I use Power BI for visualisation and reporting.
However it's good for usual Doctors offices. It's terrible for Radiology Planning. Also many institutions just put a link to their homepage or phone number in there. This way they're on DL but don't have to deal with the calendar.
...
> But maybe those smaller companies don't need to dogmatically review the checkin for a documentation fix.
It's not dogma, it's just the necessities of large groups of people working together. A small organization can use common sense, a function that scales up to (20? 50?). 10,000 people can't operate on common sense; they need another function: rules.
Much of our professional habits are part of the corporate chains which is optimised to deliver and squeeze as much as possible.
Software developed in the wild does not have those corporate obligations and the sole purpose is to enjoy the process, the sheer joy of creating something. Of programming as a creative medium.
You don't get your paintings code reviewed. It's just that artistry. You like it, then you like it, end of the story, you're not playing for the gallery.
Corporate enslavement works differently. It has moved and distributed the part of the factory shift in charge to the dude sitting next to you, cleverly. Many are just complaining to make sure they're considered the quality sensitive cooperate loyals.
You two might like different pigments for the grass and he'll strike down your painting with a red ballpoint if not to his taste.
Happens to all of us and if not, wait for it.
Declaring a single boolean flag in a corporate environment might cost more then an hour to get to a consensus because one I-am-dffierent-I-care-too-much guy has some objection about some ambiguity in the flag name in some far future and has now swayed roughly half the team on his side.
That doesn't exist in open source. Open source is all about anti status quo. It is pure rebillion. It started that way, it is about hippies and naysayers. The very root of the GNU toolchain, Herd etc are probably there.
EDIT: Typos + corporate software development environment.
seek forgiveness rather than permission - gets you launched - gets you better
That's the main value of automated tests.
Lessons learned from 15 years of SumatraPDF, an open source Windows app (2021) - https://news.ycombinator.com/item?id=35065785 - Mar 2023 (173 comments)
Seems like the author didn't look at cross-platform toolkits since 90s.
That's fair in the sense that I did not look closely at latest Gtk or Qt or WxWidgets.
That being said they certainly did not get lighter.
Last I checked Qt was over 10 MB of libraries). Sumatra is 12 MB and I'm guessing over 8 MB is fonts needed to render PDF documents.
So just Gtk or Qt code would be more than the whole app.
Latest Gtk4 does seem to look nice so maybe calling it ugly was uncalled for.
Edit: I just checked and Acrobat Reader requires 450/900/380 MB for Win32/Win64/Mac respectively [1]. One might argue that it does more than just read PDFs. But in many cases, reading PDFs is all that I need.
Discarding it out of hand by "qt is bloated" just feels disingenuous to me. You can add to that the fact that qml on the desktop if finally maturing into a viable alternative for widgets, and UI development with qml is such a breath of fresh air.
From a user perspective I remember that being a big thing some time ago when people didn't have anything already using Qt installed did an “apt install” or “yum install” on something that did and saw the small tool they were wanting was going to drag half a desktop environment in with it as dependencies. The same could likely be said for GTK in reverse, I'm not sure what their relative sizes for similar features are these days.
Some use bloated to mean the memory footprint. IIRC GTK has more of a reputation for eating RAM than Qt, but again maybe people notice a single Qt app using a lot of resource (that would be shared if running multiple apps against the same libs) when it is the only one they run.
As you suggest, just stating that “<whatever> is bloated” without reference to some details of what is meant by that, sounds a bit like someone parroting old information and/or group-think rather than having looked into it recently.
Having said that the author has a minimal dependency stance in order to try to maintain a small footprint for the app (“I avoid unnecessary abstractions.” in the section about keeping things small) so any framework that isn't little more than a cosmetic wrapper could legitimately be called more bloated than using nothing at all and talking more directly to the standard OS libs. Also in context (discussing why the product is not cross-platform and is never likely to be) this is not the only reason being given and probably not the most significant one (there may be significant selection of cross-platform issues beyond the UI framework).
The key to a lot of what is in that document is the “It’s my project and I act like it” part. All too often we forget this very important side of things, especially with one-man or small-team projects, and people comment on project decisions as if using the product gives some automatic expectation that the creator will mould it around the needs/wants of a given user or someone's idea of “the community”. For an open source project the community has the option of forking the project or offering to fund the changes they want that aren't otherwise on the creator's roadmap (though obviously the larger the project, the less practical these options may be)…
Sure, but not accessibility or UX design. There's more to a good UI than looking pleasant. Far more. I wish we hadn't unlearned that in the past decade, because Flutter and browser-based UI kits throw all of that out the window.
That said, has there been any fundamentally groundbreaking crossplatform classic UI toolkits released since the 90s? (IMGui is the only one that has seemed interesting but that's specialized and not a general one really)
EDIT: Another commenter suggest only office as example. IIRC, they are building it using chromium as front-end and .NET as backend.
I'm pretty sure that something based on IMGui could be more or less isomorphous to React style rendering, it'd be up to someone to implement it though (and it'd probably be worth it since building applications at scale you win back a lot of time by not fiddling with state manually all over the place).
Using Chromium however is explicitly not where we should want to go (it's basically a kitchen sink in itself), but we go there anyhow (me included often) because it's just so much more quick thanks to progress in dev experience in the webdev area.
I know about sioyek and been meaning to steal some of their ideas.
Maybe in the next version.
Wikipedia does the same thing, but I assume you would know that. What am I misunderstanding?
Also, if it makes your PDF reader execute arbitrary remote code, isn't that a serious risk?
And that is indeed very useful!
What the hell is Adobe doing, I wonder, that makes their software so unbearably slow and painful to use?
Probably supporting 100% of the PDF spec. plus addressing all those obscure feature requests that 6 companies in this one very niche industry really really need. Sumatra is fantastic and basically the only PDF reader I use on Windows, but it does have maybe 10% of the features Adobe acrobat has. It is however the 10% that that basically everybody needs.
I agree that Adobe has more features than SumatraPDF.
Not necessarily the PDF spec itself - SumatraPDF displays pretty much any PDF you throw at it, just way more options and stuff.
And it's not exactly slow. I'm sure the core rendering of PDF pages is same or faster.
It's just slow to start. Like very, very, sluggish slow. And it's very visible to users.
I don't think it's the features that cause the slow startup. They just don't seem to care about optimizing it.
Chrome has more features that Adobe Reader. It has video calling, a capable PDF viewer and all the other stuff in ever growing web standards.
And yet it starts up fast. Not instant as SumatraPDF but way, way faster than Adobe Reader.
I think that it's more than fair benchmark regarding complexity of the app.
The difference is that Chrome team cares about performance, including startup speed, and they spend a lot of resources on it.
I remember Chrome was counting and removing C++ static initializers from their code (the code that runs before main()) because that contributes to startup speed.
That's the level of care you need to have and I think Adobe just doesn't have it.
Sumatra includes a 3D renderer to display embedded CAD models? https://helpx.adobe.com/acrobat/using/displaying-3d-models-p...
As for startup speed, it's not really an issue as Adobe just lurks there all day in the taskbar for me at least ready to roll.
More / better annotation editing, signing, redaction is something I want to do.
OCR - I don't have a good handle on how important that is.
I would much rather throw this money your way, but I need those features
I often use PDF applications for documents that I want to keep for decades, including annotations I make. How much can I count on SumatraPDF, or any PDF application, outputting future compatible documents from conversions, annotations, deleting/merging/etc, content editing, etc.? Is there a difference between applications?
My instinct is to play it safe and use Adobe, figuring whatever they do is the de facto standard. But I strongly dislike the applications and all the privacy invasions they impose. (Yes, I'm aware of PDF/A; I'm talking about applications' outputs and not the standards.)
I will say that on older systems, Reader/Acrobat are not just slow at startup. I am writing this from a machine that has an i7-2600 and 16GB of DDR3 RAM. Reader is almost unusably slow. It's absurd.
Now where's that Mac port? /s
I’m a high school math teacher and scan dozens of textbooks every year. Adjusting a few words before printing to match what we did in class is a huge time saver for me.
Somehow my school division was able to buy me a one time fee perpetual license. I’m very happy with it.
It looks like it's Linux only (I only have Linux so I hadn't checked before), but when I want to sit down and properly read something, the keyboard-centric UI and minimalism make it a really smooth and frictionless experience.
There's a mupdf windows build. Perhaps that will do what you want.
(Mupdf is also a library so there are a lot of different programs with the name in it by various programmers.)
Maybe they're handling every posts feature wish list?
Like every comment so far here is Sumatra is nice but...< Some random feature > is missing
In fairness, I didn't write the PDF rendering. That is indeed quite a tall order.
I used to use poppler and switched to mupdf (was more active at the time, poppler seems to have picked up pace since).
The core PDF feature set isn't that bad to implement.
From what I've seen, the bad / complex parts are:
- stuff they added years later, like some XML stuff (of course you had to add XML in 2000!), JavaScript in forms
- some more complex vector graphics features like masking with vectors, support for bunch of color spaces, cmyk separation
- font handling, text rendering is surprisingly complex
- rendering fast even if PDF was badly created
- PDF is easy to screw up when you create it and boy, do people screw up in every imaginable way. You can't just say "it's bad PDF" when Adobe or Chrome opens it so a lot of effort by mupdf devs is adding heuristics to show even broken PDF docs
https://tcpdf.org/examples/example_014/
How does this even work?
You can also do animated page transitions like PowerPoint, but I don't have the right link available...
To me this is the kind of scope creep that causes a "simple" PDF renderer department to end up with 50 devs working in it full-time.
I wonder... can a PDF have an iframe that opens itself? Causing an infinie loop of it loading itself?
1. Yes we all know how stupid this was. but they wanted fillable forms and validating those forms made sense at the time. Really it was because they were trying to compete with the web.
The original spec is a bit ugly because they were saving bytes in the format (using things like single letters for dictionary keys), but some things are actually quite well thought out (Appearance Streams are great for forward and backward compatibility and are probably no. 1 reason why nothing managed to replace PDF.)
I've been wondering why they took that approach. Do you know their original reasoning? And how does it help compatibility?
[1] https://github.com/sumatrapdfreader/sumatrapdf/releases/tag/...
[2] https://github.com/sumatrapdfreader/sumatrapdf/issues/3672
MainWindowBackground = #191919
FixedPageUI [
TextColor = #282828
BackgroundColor = #ebdbb2
SelectionColor = #2d938f
...
]
what I have been doing so far is switch between other modes with autohotkey by overwriting `SumatraPDF-settings.txt`. I'd share the little script but it suddenly broke a while backNot yet but I'm thinking about how to improve it. Maybe next version.
Does anyone know of any viewers on Mac or Linux that provide these two features? Skim on Mac implements Option + Scroll and Left Click + Drag Pan, but it's not reconfigurable to any other keys or mouse buttons.
One feature I absolutely love is that Page Down goes to the top of the next page. It's very practical when you want to skim something quickly, with a zoom level that doesn't fit a page size perfectly.
Close enough?
I assume if you REALLY want to go nuclear on it, there is some shareware app that will let you do per-app keyboard emulation and rebind inputs "in flight" or something.
I believe https://www.keyboardmaestro.com/main/ is the standard solution.
It actually runs better under wine, with all sorts of errors that pop up because its updater Service can’t be found.
You can use Karabiner Elements + BetterTouchTool to rebind that when Skim is in the foreground?
The evince feature I can't live without is the find, which shows a side panel with all matches in the document along with a bit of context. I wish all document find everywhere did this.
*Which you are most likely using, since it integrates with Sumatra out of the box. And with MuPDF. And Skim, etc. It's cool like that.
I tried Sumatra a while ago, before switching over to Linux. It seemed pretty decent, nice and snappy.
I started using Sumatra years ago exclusively because of that feature. Exports/compiles failing, just because the PDF is still open somewhere is just unbelievably annoying.
[1] https://superuser.com/questions/1365482/how-to-view-a-pdf-wi...
What I meant by seamless is that all of the open source software that currently able to edit PDF document is dong it in a clunky way at best for example Krita and LibreOffice Draw. The resulting edited output document also is also looks distinct, in a bad way, from the original document unlike the output from Adobe PDF editor.
>This locally hosted web application started as a 100% ChatGPT-made application and has evolved to include a wide range of features to handle all your PDF needs
[1] https://unix.stackexchange.com/questions/85873/how-can-i-add...
Even worse PDF is not "open source", it should have been accompanied by LaTeX source.
To be fair, most PDF readers struggle with this, and I suspect only Acrobat attempts to cache or index files (as a Pro feature).
Haven't personally compared the options but for me sumatra zathura and sioyek all feel fast enough to not notice any problems.
https://ahrm.github.io/jekyll/update/2022/09/11/pdf-viewer-t...
"Now I must admit, the reason sioyek is so fast is because it creates a search index when you open the document."
Disclosure: I’m the developer behind it
When we need to sign or view some complex documents then I'll use Acrobat.
RememberOpenedFiles = false
- RememberOpenedFiles = true
- RememberSettingForEachFile = true
- RestoreSession = false
So Sumatra will open the new file without the ballast of the past. But if you click on an old file, it will open and go to the last position for you to continue. Right now one can not change the number of recent files list, nor can one replace the frequent access list with recent list, but I can live with that. Hopefully somebody could make these options.
It used to be Repligo, but it got removed from the Play store ages ago. I still had access to it, but since it wasn't being maintained, it became harder to use with each android os update.
Don't use Android 6, unless completely offline, or on a local-only network.
But I wish that I was able to ”drag and open in new window” tabs of different opened PDF files.
But I must say that skim for Mac is the best!
If so, you can do in latest update [1]
[1] https://github.com/sumatrapdfreader/sumatrapdf/releases/tag/...
SumatraPDF does support other file formats, which is nice.
[1] https://www.mozilla.org/en-US/firefox/features/pdf-editor/
It has dual page view and the option for odd or even page split - addressing your first page left/right concerns.
it is available in their Tools menu (the >> icon)
Sumatra makes heavy use of the win32 API so a port is not possible. This design is why it's so good (light, fast, well integrated) and the author refused to lessen his software by using a cross platform framework.
I will put up with less functions for the better feel.
However many people want to appeal to more users or program in Linux or web and so have a default cross platform GUI anyway (or make it cross platform easily)
If you want something still more lightweight, Zathura would be the way to go. It can be a tad too minimalistic, but great for LaTeX with vimtex, and for an occasional quick document, the startup times are phenomenal.
But one thing I really hate with it is that if you rename a file while it's open, it'll stop displaying it. Not sure if this is related to zathura itself or to the backend I'm using (mupdf).
I've been using Sumatra for past 10+ years.
It is usually the first thing on my ninite.com checklist that I use when installing fresh Windows on friends and relative's computers.
PS As a bonus Sumatra works quite well on epubs.
SumatraPDF: multi-format reader for WindowsThe author was one of the most active members of the old Business Of Software forums, IIRC.
I always wondered how he could make money from a free PDF viewer.
Works very well under WINE on FreeBSD and does not have display issues with taskbars etc.
In my opinion, I prefer the Sumatra way of doing things. But if you need to transform data... there is no other option that fighting against the multiple features of Calibre.
A nugget Bezos seems to have forgotten.
Yes, printing is not great. Actively considering making it better, maybe in next version.
In contrast the SumatraPDF dev seems focused mostly on light, casual PDF viewing. Read short pdf files from start to end, add a few highlights, occasionally search.
One good feature SumatraPDF used to have, but removed, was an option to store annotations (such as colored text highlighting) in external .smx sidecar files. Similar to how VLC and other video players support sidecar .srt subtitle files.
Since a few years SumatraPDF now only writes annotations into the PDF file. That has some advantages (greater cross reader annotation compatibility, all in one file) but also disadvantages:
- Third party tools can't search or operate on the plaintext annotations. With the sidecar approach if highlighted text were to be stored as plaintext with location data in the sidecar file third party tools could use ripgrep or similar to do fast, powerful searching and filtering on highlights across a whole pdf collection.
- Annotation versioning.
- No convenient way to load different sets of annotations for a single file. Useful for academic text work where one source document may be highlighted differently for different research topics. With an API and support for third party tools toggling such sets could be instant and very user friendly. Without sidecar files the user has to keep separate copies of the pdf file and write different annotations to each.
- Modifying the source pdf can be an issue where text integrity is crucial. Many pdf documents are used as a shared source of truth and simple hashing can verify that all parties use the same unmodified document. That gets more complicated when annotations are written into the file, since different pdf applications do such writing differently.
Even though Sioyek is better for academic work I think there is a lot more to do. In some ways academic text work on a physical book still has a better UX. For example a glance at the edge of a physical book gives an overview of the earmark locations. A digital version of that would be a code editor minimap style sidebar that overview highlight and bookmark locations in the whole pdf document. The minimap could have a heatmap feature that colors (and gradually fades) locations the user recently has viewed in the pdf. As a way to mitigate getting lost when jumping between locations in a digital document.
Am I the only person who has Preview.app constantly hang and/or freeze my Mac, even on the smallest of PDFs (< 3 slides)
I’m on M1 with latest macOS.
open - edit - reload - edit - reload - edit - Preview crashes - reopen - edit - reloads - edit - Preview crashes etc.
[ducks for cover]
https://forum.sumatrapdfreader.org/t/display-scrollbar-in-si...
I'd also find it confusing, because it seems like that means if you zoom in, the scrollbar would be for the current page, but then if you zoom out, it would be for the entire document. Is that how other viewers actually work, or is it more elaborate than that?
Obviously this isn't a model interaction with an open source project either way. Although I don't think how they responded was ideal, I don't think I would've enjoyed communicating with you on that thread if I were the maintainer, either. Nobody's perfect, but all in all, I'd hope you agree that this has thusfar not been very productive on either side.