Implementing form filling and accessibility in the Firefox PDF viewer
hacks.mozilla.org
hacks.mozilla.org
But PDFJS folks, how can you do such a subpar job at documentation?
There is really no documentation available on pdfjs. Or a proper changelog. Like in this article they talk about JS execution in pdfjs. And how they have a solution called quickjs. _no_ documentation whatsoever about how to use it.
This has been their documentation for years:
https://mozilla.github.io/pdf.js/api/
Had I created something as complex and magnificent as pdfjs I would've documented _the hell_ out of it.
I get that pdfjs usage in normal web is not their priority and what they care about is the Firefox pdf viewer but I'm sure they would've had a much better contribution rate had they done better api docs.
See the article: https://docs.lowdefy.com/generate-pdf-document-from-data
(edit: confused pdfJs for jsPdf)
Perhaps the next step is to build a Lowdefy block that uses pdfJs to render pdfs :)
It's not just a matter of documentation, but API design. Beyond the bare minimum functionality, you'll need bits and pieces of semi-public API.
For example, if you want to use the "text layer" (for selectable text) you'll need a CSS file which is inside the `web/` directory. But what's in `web/` is incomplete and not officially supported.
https://github.com/justinfagnani/pdf-viewer-element
On the docs front, I couldn't agree more. It's tough to use.
This isn't an excuse for the lack of docs for pdfjs itself, but fwiw quickjs is a 3rd-party project that's documented here https://bellard.org/quickjs/quickjs.html
I kinda miss the early 90s when we actually had competing software "houses" with some financial muscle.
Nowadays though, I don't really know since Microsoft has successfully overturned that case (probably because of Adobe being Adobe?).
Forms were appealing to some businesses and I dont have insights on that part of it, other than noting that odd media formats and bolt-on web'by things were being added to PDF at the time, and as for common forms solutions, apparently fifteen years later its still not settled.
On mobile, I use MuPDF but its limited on features and I use because it's quick.
(But upon writing that I'm not sure where I got that impression from, and I can't find a citation for it right now, so treat the claim with skepticism.)
the Quartz compositor itself has "use[d] pdf internally"[2][3] ever since the initial launch of Mac OS X
> By the time OS X came out, these were less the kind of features that "the next great personal computer operating system" should have, and more the kind that they all had. What was new included how Quartz, the engine for OS X's "killer graphics," was based on PDF.
> "What does that mean?" asked Jobs. "You know when you go to the web and you see PDF documents? Well, that technology is now at the core of Mac OS X's graphics. So you can image PDFs instantly. So now all applications get this for free."
> It's how, to this day, every Mac app can produce a PDF document without needing a third-party app.[4]
[0]: https://web.archive.org/web/20100531072734/http://www.apple....
[1]: https://en.wikipedia.org/wiki/LaserWriter#Apple's_developmen...
[2]: https://www.youtube.com/watch?v=Ko4V3G4NqII / https://www.youtube.com/watch?v=6-fkYFV7rOY
[3]: https://en.wikipedia.org/wiki/Quartz_(graphics_layer)#Use_of...
[4]: https://appleinsider.com/articles/21/03/24/apple-launched-it...
Edit: On Windows. I envy Preview in MacOS
Preview.app is king.
In other news, that's precisely why I'm weary to open PDFs in (edit: recent versions of) macOS because it autosaves, and I need to have a bit-perfect copy of said PDF. I mean, I open it in Firefox or Chrome, and this is just a problem that is specific to my circumstance, but it still bugs me.
They abandoned Mortar when Adobe have announced that it will retire Flash.
(Also, PDF.js is now a misnomer: it now mainly uses WebAssembly.)
--------------------------------------------------------------------------------
Language Files Lines Blank Comment Code
--------------------------------------------------------------------------------
JavaScript 355 144428 13069 16759 114600
JSON 12 24385 2 0 24383
CSS 12 3302 407 183 2712
HTML 32 1638 176 230 1232
Markdown 19 733 221 0 512
TypeScript 1 20 4 2 14
CoffeeScript 1 15 2 0 13
YAML 1 4 0 1 3
--------------------------------------------------------------------------------
Total 433 174525 13881 17175 143469
--------------------------------------------------------------------------------[1] https://poppler.freedesktop.org/
[2] https://gitlab.freedesktop.org/poppler/poppler/-/issues?labe...
[3] https://gitlab.freedesktop.org/poppler/poppler/-/issues/463
[4] https://gitlab.freedesktop.org/poppler/poppler/-/issues/230
[5] https://gitlab.freedesktop.org/poppler/poppler/-/issues/364
There are two disturbing things in that sentence. At least the JavaScript I can see a reason for, it can check whether your input is correct and maybe autofill or things like that, and should of course be sandboxed. The telemetry part is more scary.
If Firefox occasionally showed me a report on some metrics it collected and asked me if I was okay sending it to Mozilla, I'd probably be fine with it. However, the choice to take it without my consent will always be opposed.
It really feels like a forced way to try and shoe-horn something good to say about telemetry when it is something everyone that deals with PDFs have known since forever.
And if that is the best they have to show for it ...
A Web crawl could tell you how many PDF documents on the Web use forms, but it won't tell you often your users encounter those documents --- many documents aren't on the public Web, and some documents will be encountered far more often than others.
the way the world has done it forever... a survey?
it is concerning. if Firefox is doing it, they are all doing it. i am not going to be opening PDFs in browsers.
this is ridiculous.
With better data you can make improvements that help more actual users instead of just focusing on what vocal minorities think.
Now what you might miss is that in some country where special circumstances led to an unusual high use of a certain feature.
But this is not that case.
More importantly, even with telemetry. Is the JavaScript adding something or is it solely used for tracking? Is it appreciated by the users?
Maybe 80% of input validation is so broken that they only serve to block legitimate input and waste cpu-cycles.
More than telemetry you need common sense. It is not apparent how the telemetry influenced this in any way.
Not everyone needs to care about everything all the time. I would complain about JS in PDF too, but I don’t have a voice anyway.
There's a nice big section in the Firefox settings called "Firefox Data Collection and Use" that has some fairly fine-grained options for what data you feed back to Firefox. I have all of mine unchecked, but I could imagine that if I were more community-minded, or if I were interested in actively contributing to FF's development day-to-day (for example running on nightly builds) I would activate those.
FWIW, I do explicitly allow telemetry for the applications I buy licenses for and use for my work; it's in my best interest for those programs to be improved as efficiently as possible (and to have my use patterns be part of the corpus used to prioritize those improvements), and it makes it easier for me to report issues.
I looked into these settings but I can only find 4 checkboxes. Which of the four checkboxes would this telemetry fall under? Is this a study, the result of crash report logging, or is it "data about your interactions with Firefox [..] (such as number of open tabs and windows; number of webpages visited; number and type of installed Firefox Add-ons; and session length) and Firefox features offered by Mozilla or our partners (such as interaction with Firefox search features and search partner referrals)"
I know Firefox collects some technical tracking, but there is no up to date overview of what tracking is actually done. Is there a list of parameters tracked like Microsoft published [1] after the outcry about their terrible tracking? The privacy policy is vague and the documentation for developers [2] contains phrases such as "opaque prio-specific payload. Like { a: <base64 string>, b: <base64 string> }" which isn't exactly useful.
[1]: https://docs.microsoft.com/en-us/windows/privacy/required-wi...
[2]: https://firefox-source-docs.mozilla.org/toolkit/components/t...
I don't see the same screen you do (perhaps something to do with telemetry being force-enabled in nightly?) so this is based off of what you wrote:
> Is this a study...
Studies are generally minor tweaks in user configuration sent to a subset of users (often nightly), and then telemetry between control and experiment groups is compared. You can see what studies you have been in at about:studies. For example, some of mine include setting the user-agent to Firefox 100 (to test a 3-digit number), enabling fission (and a number of things with process counts and whatnot), enabling HTTP/3, and (I kid not) "Changing a pref that does nothing - to check enrollment and unenrollment reliability." A lot of the time if you inspect element the source has a bug number you can look at (which really ought to be a built-in feature).
> The result of crash logging
You can see this in about:crashes; this is when the browser or a process (eg. GPU) crashes.
> Data about your interactions with Firefox
Probably, this seems like the most broad and undefined one
> Number and type of installed addons
Seems self-explanatory
> Firefox features offered by Mozilla or our partners
Not really sure about this one, presumably it relates more to things that have to do with third-parties or them making money
> Overview of tracking
See about:telemetry. That has links to a number of other pages I didn't realize existed before today, but which seem to have a pretty comprehensive public overview.
about:telemetry didn't work on my phone (although the screenshots and videos I could find on it seem to indicate that it's just names and raw values, not human-readable descriptions), but it did contain a link to https://probes.telemetry.mozilla.org/ and the newer https://dictionary.telemetry.mozilla.org/ that seem to contain an overview of telemetry data collected.
I don't see anything on there about PDFs and Javascript, so I'm guessing there's another telemetry endpoint out there that I'm not seeing. Maybe the PDF renderer has a separate telemetry system, I don't know. Seeing the amount of (IMO useless) data Mozilla collects, I'll be opting out from telemetry from now on anyway.
Makes sense. How do you see Firefox as different; you don't pay for it, but, if you use it, you presumably still might want it to be improved efficiently and based on your use patterns. But no?
B) The tools I use are force multipliers for me, that differentiate from their competition through the improvements they make (thus making me far more productive than I otherwise would be). I use Firefox, but there isn't a huge amount to distinguish it from other browsers, and it's not clear to me that improvements in browsers over the last 10 years (beyond security) have added enough value to my life to need my data.
C) For the tools I'm using, I am one of a relatively small number (tens of thousands) of users, I am occasionally in direct contact with their staff (again, I pay for this support, among other things); my contribution of telemetry data to them is making a substantial contribution. For Firefox, with its install base in the hundreds of millions, my individual contribution is meaningless.
The fact they had to find via telemetry something that every developer who had interacted with a PDF with forms knew (that they have JS) is very worrying.
I'm expecting a long long list of bug reports when they start finding the mountains of corner cases not shown by their "telemetry".
Lately as I've been working on mapping projects and want to view large PDFs I've found many which need to be saved and opened in Preview.app instead of the browser.
This is a great example of a map which views horribly in Firefox (try zooming in) but works great in Preview.app: https://www.bia.gov/sites/bia.gov/files/assets/public/webtea...
I think a PDF viewer (I prefer "reader", because that's what Adobe called their free product) is for viewing PDFs. The point of PDF, to my mind, was that you got a much higher level of control of typography and design than you could get with HTML. So designers and marketers liked it. I have never bought into interactive PDFs.
(I don't know details of this specific case, and whether/why the plans and decisions made sense at any given time.)
> Mozilla developers at the time actually have other plans with the Pepper plugin support: Adobe Flash (since that Chrome's Pepper Plugin API (PPAPI) is actually miles better and secure than the Netscape [-era - ed.] Plugin API or NPAPI). They abandoned Mortar (the name of the project) when Adobe have announced that it will retire Flash.
It's because pdf.js does not print vector and instead is just printing canvas bitmaps. If they ever get the SVG backend working its will finally have good print quality without trying to get OOM doing measly 300 dpi canvas.
There is a lot of open source software out there that probably already can be compiled to run in a browser with WASM.
I mean it wouldn't really be "ads", more like suggested extra pages provided by trusted partners...
Residential mortgage application? "Shop Rates and Save"