CVE-2024-4367 – Arbitrary JavaScript execution in PDF.js
codeanlabs.com
codeanlabs.com
https://stackoverflow.com/questions/49299000/what-are-the-se...
Not only does it correctly identify the attack vector of this CVE, but I think his advice on how to mitigate it is sound. Is there something I'm missing? The only flaw I see is that it doesn't consider the implications of using PDF.js in Electron.
The option isn't supposed to allow XSS-by-design (which the original requester was worried about), the possibility of a vulnerability is mentioned, the impact of a vulnerability is correctly described (XSS not RCE or similar), and mitigations that would effectively limit the impact of such a vulnerability are presented (separate origin).
"Potentially, there is a tiny possibility... Keep in mind, it's a web application, and worse it can do is XSS attack"
One sad thing about security is that everyone and their uncle is a security expert.
People who have been in this game more than six months would never making such a claim.
And only XSS? What does that mean in the context of the page, or an electron app? How can this guy know "just an XSS" is not catastrophic?
First off, are we not supposed to have "random guys" writing stuff on Stack Overflow and Wikipedia? Because that's kind of how those websites work: they rely on "random guys" to do all of the writing, rather than relying on credentialed experts only. I sure think Stack Overflow and Wikipedia are very useful resources despite having "random guys" do all the writing.
Secondly, you attack the random guy for... correctly identifying that "the worst it can do is an XSS attack". This is very useful and accurate information. Information like this is typically absent from all kinds of vulnerability disclosures. When you read on the news that something something has a vulnerability, they typically they don't give you the practically useful bit of information, like what is the practical scope. Is it a 0-click RCE or is it a XSS inside a web app? They don't tell you. Except this random guy, who accurately identifies this information.
> How can this guy know "just an XSS" is not catastrophic?
"Just an XSS" is the correct description of the severity here.
More like dogpiling and coattail-riding of the current in-focus topic. Both comments smack of smug know-betterness but are accompanied only by vague remarks and no real claims that might be subjected to scrutiny. It's almost like dogwhistling for karma.
"If you serve untrusted PDFs in a PDF viewer and you are hosting at your location, it is better be located at different origin than your main app, www.example.org vs pdfviewer.example.org."
My understanding is that this CVE is an XSS attack, so isn't this advice sound? The RCE portion of this CVE is for Election, where every XSS attack is twice as fun.
Is there something about his answer that is wrong that I don't see? I hardly think we can fault someone simply for having faith in the integrity of software that everyone else trusted until now.
Is the origin right? Are all the security headers correctly set? Is it even possible to keep up with all the stuff that is published, today, to sort of try and secure a web app? I don't think so.
My approach is... no JavaScript (script-src 'none'). Just don't do it.
You realize this particular problem requires JS to solve, right? It's not just there arbitrarily.
That's like saying your approach to web application hardening is to disable all inbound connections. You quite necessarily need those for a web application.
I can't see how this would be pragmatic or productive, or maybe it's not meant to be.
Most of the time, you get better results by just serving up the PDF as Content-Disposition: inline. 90% of JavaScript PDF viewers are crap (although PDF.js is decent), but the browser on all dignified platforms still does better than any fancy JavaScript I’ve ever seen. Loads faster, too.
I don't think anyone but a bored hobbyist is implementing or using a JS PDF renderer solely for rendering and the people who are using it for the aforementioned reasons don't care about load times, the document has to be processed somehow by the user agent.
The suggestion that JS is optional here is nonsense and that the built in browser PDF renderer (written in a lower level language than JS) is faster is common knowledge.
Again, I don't see how this suggestion is pragmatic or productive, but still, maybe we're not trying to be here.
You also can prepopulate the form with user input from previous webpages/the user account. This is how a lot of HRISes do it.
Usually people want onboarding to be as frictionless as possible. Downloading the PDF and using some editor outside the website then coming back to upload it counts as friction.
And this conversation is still far away from the point: you need JS for a JS based PDF renderer, and there are valid uses cases where one is required.
> Usually people want onboarding to be as frictionless as possible. Downloading the PDF and using some editor outside the website then coming back to upload it counts as friction.
Wait, are you saying there's a workflow for which the most frictionless solution is to have the user fill out a PDF, on an authenticated website and submit the PDF in the browser? As opposed to, say, <form>? Can you elaborate?
It would have worked much better as an emailed PDF, or a simpler form. Or an ordinary non-PDF form that would generate a filled-in PDF that the user could then sign.
> I'm going to discontinue this conversation though.
Oh, well.
I'm sure a lot of people here will disagree but most websites don't target the minority that prefers an email based workflow.
I tend to agree outside the context of a browser, but the post is about PDF.js.
> fill forms in the browser (where their authenticated context resides)
I have never, in my entire life, seen a web form implemented as a PDF where this was anything other than miserable. Use a form, thank you very much. The less SPA magic and the fewer intermediate submit buttons, the better. (Maybe, if the actual goal is for a user to fill out a form and print the result, offer PDF.js to users on Windows who might otherwise be stuck with Acrobat. I distinctly remember Acrobat being the best PDF-form-filling software maybe 20 years ago, but Acrobat has gotten, if anything, worse, and basically every other package out there runs circles around it.)
Frankly I'd just disable script evaluation if you don't specifically need that.
And how do you know whether you "specifically need that"? As the answer says, it's not for scripting within the pdf itself, it's for optimizing font rendering. For pdfs that you don't control, it's basically impossible to know whether that'd be needed or not. Even for pdfs that you do control, in a large company it's very likely that the team that's configuring pdf.js isn't talking to the team that generates the pdfs, which means you have a similar problem.
This vuln works even with scripting in PDF.js disabled.
> PDF.js runs under the origin resource://pdf.js. This prevents access to local files, but it is slightly more privileged in other aspects. For example, it is possible to invoke a file download (through a dialog), even to “download” arbitrary file:// URLs. Additionally, the real path of the opened PDF file is stored in window.PDFViewerApplication.url, allowing an attacker to spy on people opening a PDF file, learning not just when they open the file and what they’re doing with it, but also where the file is located on their machine.
If you're just viewing a PDF using the built-in pdf.js in Firefox, then (AFAIK) it doesn't matter what site you downloaded it from, because pdf.js isn't running in the context of the website, so it doesn't have access to that site's locally-stored data (including cookies). Instead it's running in the origin mentioned above, with the accompanying concerns.
So the XSS (again, as far as a web browser is concerned) would be if the site itself is shipping pdf.js for viewing PDFs inside the webpage itself. As you suggest, Gmail lets you preview PDFs, so XSS would be a concern there, but only if Gmail is using pdf.js.
This too can be harderend against, but it's a significant attack vector in quite a few desktop applications if users don't update.
> In applications that embed PDF.js, the impact is potentially even worse. If no mitigations are in place (see below), this essentially gives an attacker an XSS primitive on the domain which includes the PDF viewer. Depending on the application this can lead to data leaks, malicious actions being performed in the name of a victim, or even a full account take-over. On Electron apps that do not properly sandbox JavaScript code, this vulnerability even leads to native code execution (!). We found this to be the case for at least one popular Electron app.
So still no chance in the foreseeable future for this monstrous "paper-based" mockery of docs in a digital age to get phased out?
(Except PDF/A-4, which reintroduces JavaScript for some horrific reason).
It was pdf.js handling of fonts
Copy paste mostly works fine for me. I only have trouble when it's generated in a weird way (eg. scanned from a paper document then fed through OCR), or has complex formatting (eg. math equations) that have no hope of working correctly in any system. In those cases, I don't see how it's the fault of the PDF format, any more than HTML (or whatever you think is a "real digital document" format) can embed a picture of a scanned document that totally breaks copy-pasting.
(but also math equations have plenty of hope even though they're complex indeed, you can copy&paste some kind of "latex" representation that is sometimes used to ... produce those PDFs)
> whatever you think is a "real digital document" format
whatever supports basic digital interaction we've had available to use for many decades in alternative formats, or whatever doesn't have those rigid pre-digital-paper-based layout limitations where you can't use one of your most popular digital devices - your phone - to read a doc since the phone is smaller than a sheet of paper
> PDF file format (which supports semantic paragraph tags, for example).
These are called newlines and have a pretty widespread support outside of some paper pockets of resistance! You only need some other semantic tags because the format fails at basics
Here is one from Adobe https://www.adobe.com/support/products/enterprise/knowledgec...
Or even better: their annual investor docs a team of professionals has spent time carefully preparing...
like this https://www.adobe.com/pdf-page.html?pdfTarget=aHR0cHM6Ly93d3...
(but don't look at the annual report, that marvel of a public disclosure document not only doesn't copy&paste paragraphs, but has another nice niche use of PDF - you get garbage chars instead of text, rather ironic)
https://www.adobe.com/pdf-page.html?pdfTarget=aHR0cHM6Ly93d3...
[1] https://www.federalreserve.gov/mediacenter/files/FOMCprescon...
The format supports a lot that is not commonly implemented by PDF readers (or PDF producers).
And a good format wouldn't require any ToUnicode maps for simple text in the first place
And poorly supporting a lot without common implementations isn't a defence against the charge of high complexity and bad design, but a reinforcement thereof
(also, no, the first document doesn't work on iOS, I select title and two paragraphs, copy, paste, and I get a single line instead of 3, so a different manifestation of the same common fail of PDFs)
Still, the fact that some PDF processors can make this work shows that the format isn’t broken “by design”.
https://ecma-international.org/wp-content/uploads/TC46-XPS-W... and https://ecma-international.org/wp-content/uploads/TC46-XPS-W... are interesting in that they're different packaging of presumably the same data for compare-and-contrast. I will say that exploring .xps files is much easier via $(unzip) than using qpdf or friends
Unfortunately, applications that produce broken PDFs are rife, and Postel's law sets the expectation that we should consume garbage and be happy.
Garbage inputs are the responsibility of the sender, not the receiver. You can and should accept a small margin of error in inputs where errors may logically appear, but if the receiver accepts too much error then it becomes responsible by creating a complicit norm. If the responsibility of error remains on the sender then introducing further validation is less likely to cause breakage in communication.
- https://www.gnucitizen.org/blog/danger-danger-danger/
- https://blog.jeremiahgrossman.com/2007/01/what-you-need-to-k...
https://en.wikipedia.org/wiki/Principle_of_least_astonishmen...
Common sense suggests PDFs are the digital equivalent of paper documents. Paper documents can't run Javascript, so PDFs shouldn't either.
> You might be surprised to hear that this bug is not related to the PDF format’s (JavaScript!) scripting functionality. Instead, it is an oversight in a specific part of the font rendering code.
It goes on to explain that pdfjs dynamically constructs and executes javascript functions as an optimization for rendering older fonts. Certain arguments pulled from the PDF were not escaped, validated, or delimited (the values were expected to be numbers), so you could inject arbitrary JS. (At least that's how I read it.)
I don't use Firefox to view PDFs. But it seems that the new MS Windows echosystem might be affected.
Why? Does the microsoft pdf viewer use pdf.js internally? Edge at least is based on chromium, and chromium AFAIK uses pdfium rather than pdf.js.
Wonder if it was related...
This made me chuckle
Hidden in some paragraph it does say
> Instead, PDF.js runs under the origin resource://pdf.js. This prevents access to local files, but it is slightly more privileged in other aspects.
Seems like it's not an XSS letting you take over the website origin, but it lets you run JS under this resource://pdf.js origin. Could be an interesting vector when combined with other weaknesses, but not an instant knock out as I expected when I read the title and saw the points :)
You are right for the case where Firefox's PDF.js is used (local or remote file in a tab or iframe). The XSS problem however is with web-applications that themselves use PDF.js. In that case, it does not run in a separate or special origin; that is a Firefox thing.
You are also right that the PDF format supports JavaScript, but that is something unrelated to this, and indeed highly sandboxed in all cases.
I suppose/hope this is version number and not a dig at the market share
With the exception of LTS releases, if you haven't got firefox 126 yet because you're on a "stable" package manager, I'd encourage you to promptly download firefox from mozilla.org (which will come with auto-updates) and uninstall your package managers insecure version. Given the state of the web and software security web browsers aren't something you should be delaying updating by a week.
Which distros have this problem? AFAIK debian-based distros (eg. debian, ubuntu) package firefox ESR which is kept up to date with security patches.
Yes, a fix landed in Firefox, but the vuln is in pdf.js, and now I’m giving the ol side-eye to the four or five electron apps I have running.
It's already fixed in Debian stable (firefox-esr version 115).
It's fixed by default in FF 126+. But, as I understand it, older versions like the one in Debian stable, can be (and are already) patched.
$ sudo snap refresh firefox
firefox 126.0-2 from Mozilla refreshed
I'm not always happy about the snap mechanism, but this time I'm glad about a quick release/packaging channel. Kudos to the firefox snap maintainers!I don't see the relevance of snap here.