Insecure Features in PDFs (2021)
web-in-security.blogspot.com
web-in-security.blogspot.com
I wonder how one would measure the costs and benefits (with a focus on security) of speeding up and making more security-driven the gargantuan task of shifting to a better, well, portable document format. People are thinking big and want measurement, eg just posted https://news.ycombinator.com/item?id=39514844 WH release on memory safety https://www.whitehouse.gov/oncd/briefing-room/2024/02/26/pre... would it not make sense to be similarly ambitious and metrics-driven for this too?
pdftotext file.pdf - | nroff | less
pdftotext is from Poppler Developers http://poppler.freedesktop.org Glyph & Cog, LLC
nroff is a GNU common util on most linux/unix systems
(though I don't trust poppler utils to be secure)
Edit: re-reading I guess your point is that you can use tools to extract text from PDFs and then read without worry. That brings up another super annoying thing about PDFs--it's can be hard to extract text from them with high fidelity.
Or impossible with scans. Usually I take a look at the filesize as a guestimate. Anyone know better FOSS OCR tools that run from cmdline?
In practice, it's a simple JS kaboom that they didn't catch, because error handling is for n00bs
You can embed a port scanner in JavaScript inside a PDF document via /JavaScript tags.
And many high-end IDS/IPS/NDS/XNS firewalls look into PDF documents for such things like that.
https://opensource.adobe.com/dc-acrobat-sdk-docs/library/sam...
That being said, in combination with exploitable vulnerabilities in particular PostScript parsers used in software like Acrobat, the Turing-completeness of PostScript makes it much harder to detect such exploits. 0days in PDF readers are "nice exploits to have", because a PDF virus can be coded such that it's represented in code inside the PDF in an unbounded number of different ways, foiling signature-based virus sanners.
gs -dNOSAFER zmachine.ps -- yourgame.z3Thankfully we learned from them and didn’t repeat that mistake over and over again.
I’m not a PostScript expert but I’ve been reading a lot about it recently. It’s a rather fascinating system for 2D graphics.
1. a declarative language for describing the same things PostScript describes,
2. which allows the rasterization of arbitrary shapes at arbitrary DPI (PostScript is DPI-oblivious — it's up to the printer what DPI it's printing at!);
3. and which also works for vector plotters, that will never rasterize the data you're sending at all, but will actually follow the bezier curves, like the 2D version of 3D-printer GCODE;
4. and which enables the implementation of this rasterization and/or plotting on a variety of affordable hardware architectures in the 1980s — where the 1980s was a time where CPU power wasn't too expensive, but where memory prices were at an absolute premium. So your printer might have had a CPU as powerful as your computer's in it, to crunch PostScript — but definitely wouldn't have had the memory to buffer a full rasterized page.
---
To put that last constraint another way: PostScript was designed to be rasterized in a way that enabled printers to do something much akin to "Racing the Beam" (https://www.youtube.com/watch?v=sJFnWZH5FXc).
In both the display-rendering and printing cases, this was done in 1980s hardware, because 1980s memory was too expensive for most systems to be dedicating it to hold a buffer to asynchronously pre-render into and then read from when drawing.
So instead, in both cases, you must render+rasterize in one motion, programmatically and extremely efficiently. And the obvious way to do this, is by using a CPU with rasterization MMIO registers it can very quickly poke at, to change mode bits during the rendering process. A CPU whose ISA becomes, in effect, a Domain Specific Bytecode for procedurally generating raster-lines.
If printer vendors of the 1980s could have been expected to agree on a single such ISA, then chip vendors would have just made printer SoCs that conform to that ISA — and we'd have ended up with some kind of "vector-drawing abstract machine" bytecode (with real hardware impls in printer ASICs, but also virtual ones on PCs) rather than PostScript.
But as with RDBMS vendors in the 1980s, the printer vendors were all too invested in their own internal architectures to agree on what the low-level execution plan should look like.
And a with RDBMS vendors, the solution that all these 1980s manufacturers could get behind, was a standard for a (theoretically) portsable, text-based intermediate-language standard — one that could be generated by computer software, transmitted to their system, and then, inside their system, compiled down to whatever internal representation allows the system to do things its own way.
In RDBMSes, this "intermediate language" was SQL. In printers: PostScript.
It also confirms many suspicions I've had over the years that have led me to, e.g., running all PDFs from questionable sources through VirusTotal before viewing on platforms where I wouldn't normally run antivirus software.
The original article also confirms my suspicions that this step is inadequate:
Because the Launch action can be considered as a dangerous feature, we conducted a large-scale evaluation of 294,586 PDF documents downloaded from the Internet, in order to research if there are any legitimate use cases at all. Of those documents, only 532 files (0.18%) contained a Launch action. While none of the files was classified as malicious according to the VirusTotal database, we conclude that the Launch action is rarely used in the wild and its support should be removed by PDF implementations as well as the standard.
Incidentally, the Launch action is still present in the most recent version of the PDF standard[1], with only OS-specific launch parameters deprecated (which include passing arguments to the launched executable, so eliminating deprecated features is still a significant security gain).
Finally, I'm both personally and professionally curious about how the non-DoS examples in this articles may apply to non-viewer PDF tools and libraries like qpdf[2] and Ghostscript's original and recently reimplemented PDF interpreters[3].
[1] https://pdfa.org/resource/iso-32000-pdf/
(registration required, but at least the base standard is available at no cost; sadly, important incorporated standards like ISO 21757-1:2020 [ECMAScript for PDF] are not)