Converting untrusted PDFs into trusted ones: The Qubes Way (2013)
blog.invisiblethings.org
blog.invisiblethings.org
https://github.com/freedomofpress/dangerzone
> Take potentially dangerous PDFs, office documents, or images and convert them to safe PDFs.
From the learn more about page:
> Dangerzone was inspired by TrustedPDF but it works in non-Qubes operating systems, which is important, because most of the journalists I know use Macs and probably won’t be jumping to Qubes for some time.
> It uses gVisor sandboxes running in Linux containers to open dangerous documents, instead of virtual machines. And it also adds some features that TrustedPDF doesn’t have: it works with any office documents, not just PDFs; it uses optical character recognition (OCR) to make the safe PDF have a searchable text layer; and it compresses the final safe PDF.
Previously (announcement and details of gVisor sandboxing etc):
Safe Ride into the Dangerzone: Reducing Attack Surface with GVisor
Seems that PyMuPDF is used with a fixed (single pathname) "/tmp/input_file" ? https://github.com/freedomofpress/dangerzone/blob/main/dange...
Everything else is tossed through LibreOffice.
Meanwhile what I'd prefer for PDFs is some allow-listed set of 'safe' PDF operations (layout image, layout text) to be used with sanitized inputs (no underflows, overflows, corruption, etc), and the results of any at snapshot-runtime code evaluated and then flattened out to a safe element. Image OCR could be run atop that.
Similarly it'd be nice if a filter like that existed for the other documents, but as an individual contributor I don't have the human power to keep up with that goal and would take the same low hanging fruit worse but secure output route.
That's the problem right there. PDF supports many image formats, including ones that are useful but you may have never heard of like JBIG2 for scanned documents. And the parser for those image formats needed to be secure as well. One very famous exploit is just exploiting JBIG2 (among other things): https://googleprojectzero.blogspot.com/2021/12/a-deep-dive-i...
A single fixed path for the temporary input file means only one copy of the converter can run at a time (unless it's in a sandbox, which wasn't clear from the context). A PID specific tempfile or, better, use a standard 'make a new temp file name so there's not a collision' utility, would allow parallel conversions.
> Someone could potentially use the underlying tools directly in their own needs workflow, or an expert in those tools could mention if there's a vulnerability / common mistake pattern.
Yep, that's something we cover in our about page. We process the untrusted document in a gVisor sandbox, that runs within a hardened Linux container, that runs within a Docker Desktop VM (in Windows/macOS). And yet, that doesn't mean that Dangerzone can't get hacked by a determined (possibly state-backed) attacker. So yeah, you have a point.
At the same time, using tools like this raises the "you must be THIS determined to enter" bar. This means attackers must spend much more money, time, expertise, 0-days, etc. These resources are finite, so more people are protected as a net result, even if we can't protect everyone. That's the way I see it at least.
https://github.com/QubesOS/qubes-app-linux-pdf-converter
Their source code seems to take the most obvious path... flatten it to an image printout then possibly do more? https://github.com/QubesOS/qubes-app-linux-pdf-converter/blo... https://github.com/QubesOS/qubes-app-linux-pdf-converter/blo...
Though at a quick skim I can't see any OCR steps.
Two more things can happen.
The increasing volume of memory-safe utilities means they can be used on one or both sides of this. That might prevent the exploit entirely. If a memory-safe CPU, it can still help to isolate in case of hardware failures (esp bitflips).
It can also be used to boost performance in non-Qubes systems where a secure (or OSS) processor is in use. They’re often slower than commodity CPU’s. So, one can use the disposable VM’s on commodity CPU’s to filter data (block most attacks), transform it, and send it over simple, wire protocol. Commodity VM’s might also present it back to the user in dressed up form.
Outside of security, a long time ago, they were doing similar things to decrease latency and boost bandwidth on Beowulf clusters. A team made Fast (or Active?) Messages to eliminate TCP/IP as a bottleneck. So, sometimes a security technique can also be a performance booster.
I love the idea of making PDFs dumber and safer but maybe ePub would fit the bill? I'm just thinking out loud, I would like to do this again, but the Qubes way of spinning up a disposable VM to produce a monster PDF file is unsatisfying. More general Qubes being slow was a big reason I switched off of it
Converting untrusted PDFs into trusted ones: The Qubes Way (2013) - https://news.ycombinator.com/item?id=10538888 - Nov 2015 (5 comments)
Surely you can do that instead? Parse the PDFs and format them in basic ways without support for "extensions" or anything. Let the user read that before using the "real" document with extensions potentially enabled.
So there is a certain sense of absurdity of needing to spin up an entire VM just to render a PDF. Running a standard PostScript renderer in a user executable (perhaps in a chroot jail to be a little bit paranoid) should be enough for safety. Or just stick it inside a Docker.
Restrict the permissions on the user process to “read my static data files like fonts” and “write output to this 1 file, or a parsing error to this other 1 file”.
Expected by who? It's been associated with security bugs for decades.
Here, for example, is but one of the scripting engines https://helpx.adobe.com/acrobat/using/applying-actions-scrip...
Here’s its JavaScript execution engine for embedding JavaScript in a pdf: https://opensource.adobe.com/dc-acrobat-sdk-docs/library/jsa...
Your belief is what makes them an excellent exploit deployment format. “What can go wrong? I know tech and they cannot be harmful.” Click.
The author made this seem like such a fundamental issue. Is that because PDFs natively have support for say executing code (i doubt) or accessing the filesystem (i doubt), etc...
Yep, here’s an Acrobat Reader release from two days ago that fixes two arbitrary code execution vulnerabilities since the previous one two months ago: https://helpx.adobe.com/security/products/acrobat/apsb24-92....
I haven’t looked into browser-embedded PDF viewers enough to know how they compare to other software – they’re definitely much safer than Acrobat and still not completely safe (e.g. CVE-2023-1530 in Chrome wasn’t that long ago) – but I would expect them to be at least as safe as other browser functionality.
> Is that because PDFs natively have support for say executing code (i doubt)
They do (https://helpx.adobe.com/ca/acrobat/using/applying-actions-sc..., including “Run a JavaScript”, although that has to be enabled), but indeed that’s not the one fundamental issue; it’s usually just standard vulnerabilities of memory unsafety or terrible design (XML).
If I remember correctly Google bought a source code license from some Aussie company (?) for rendering PDFs in Chrome. That was like a decade ago though. I wonder what happened since. Probably lots.
https://web.archive.org/web/20140529210328/http://www.foxits...
> Founded in 2001, Foxit is a leading software provider of solutions for reading, editing, creating, organizing, and securing PDF documents. Headquartered in Fremont, CA, USA, Foxit has operations worldwide in China, Belgium, Japan, and Taiwan
seems like most of their presence was in china, and was domiciled in china, but they had sales "offices" in other countries and so they emphasized that part for better PR.