Embedded PDF viewer in Firefox 81 supports filling forms
support.mozilla.org
support.mozilla.org
We don't know whether pi is that either for any integer base.
Sure we do. There are plenty of proofs out there that pi is an irrational number.
I understand the point that PI contains every possible piece of information, theoretically.
However, the chance of finding a given string in PI depends on the string’s length. The longer the string, the more the probability tends to 0.
The paradox therefore is that PI contains every PDF, but you will never find them, so in what sense does it really contain them at all?
Including a PDF that generates the digits of pi
[1] https://www.pdflib.com/pdf-knowledge-base/pdf-password-secur...
And even places with strong roaming rights tend place limits on well marked land.
Probably won't send it to any recruiters, but it will be a funny anecdote for interviews.
I've read a few comments on HN how PDF is, well, not developer-friendly. If people are interested in providing some more examples here, I'd be curious to know!
I personally think the structural PDF format is a really great format. It's entirely ASCII-based, a pure text format, yet it can embed arbitrary binary data and compress that data. The actual structure is simple and support just enough functionality, like a tree of object, dictionaries and arrays, unicode strings, date formats, etc.
I think if you limit yourself to pure structural PDF woulde have been a great format to standardize upon, much better than JSON or XML. It';s richer than JSON, simpler and saner than XML. Again, it's top-notch ability to embed binary is great. It has other great characteristics, for example you can update anything just by appending.
The ugly bits are in the "semantic" PDF: the page descriptions, media, etc. Even then, the early version of PDF were nice, mainly just simplified Postscript.
A common case is clients who use utilities that generate single customer documents then merge them into a bigger file for bulk print and mail (bills and statements, not identical copy). Without fail that results in thousands of similar but different subset fonts whereupon most printers I've encountered eventually fail due to memory issues.
Typically this leads to a discussion about "I can open it on my computer fine" and bending over backwards to find a workaround. Merging and consolidating these fonts doesn't seem to be a simple task, although some tools claim to work some of the time.
Something that scopes object resources for disposal could be nice in the PDF spec (maybe it exists), but something like a LRU caching mechanism on the printer would potentially resolve this too.
Having said that (and worked on a commercial PDF library), despite all the cruft that came with age, it's a well built format that survived the test of time with good reasons.
I worked on a Pdf with inbuilt tracking solution that updated the form layout using ActionScript based on the workflow status and the role of the user (ie the line manager had a different group of fields in the form to complete like their signature while viewing what the initial requestor had entered) Lots of callbacks to the server saving in progress data and updating the status of who had the form, who was next based on department and emailing it to that person if it passed validation.
An initial fun discovery was that you could force the form to download and replace itself with the latest version even if they had just opened some old file they had on their pc.
Streams just suddenly end? That’s ok. Totally corrupt xref tables? Ok. Incorrect image headers? Ok. Unrecognisably mangled Type1 font formats? Fine!
A lot of the time developers want access to the text inside PDFs. Unlike HTML or formats like MS Word (XML or old binary format) getting "text" isn't really possible.
Most "document" formats have the concept of words or strings: a set of characters separated by whitespace. PDF isn't a "document" format in that sense - it's a page description language. Instead of strings of text you have character glyphs positioned at a particular location.
If you want to "read" the text, you have to work out the orientation (which can change throughout the page - think of table header alignment), and use some kind of heuristic to guess the word spacing based on the font and character spacing.
There's also this whole thing with clipping, where some text can be hidden behind other objects (or off page) so you have to try to deal with that.
There's lots of libraries that try to do this for you, but there are lots because none get it right 100% of the time...
- Clipping path logic, you can write text outside of it, which makes it effectively invisible yet it will show up if you try to extract the text.
- Anything regarding the graphicstate stack, it's a pain to debug.
- Extracting content from AcroForm/JS "XFA" forms
PDF is great format for printing, it's just a pain for pretty much everything else.
My other one is the use of multiple subset fonts that are actually the same font with a different subset of glyphs that you want to merge back together.
EDIT: oh yeah, I'm pretty sure it contains Mail, just like Zawinski said
The old .Doc and .xls files were a bad format, but my understanding is that since Office 2007 the format is generally much better.
Well, at least, pdf is probably better than printed paper for that purpose.
That is truly nightmare material. Especially considering a non-trivial percentage of pdfs circulating are non-conformant and people still expect them to render...
Mozilla's pdfjs[1] project is a pure HTML/JavaScript solution for PDF rendering. This is the same code that ships in Firefox browser as well. This is standalone, AFAIK, it doesn't talk to a mothership.
Ended up just making the app generate HTML before calling wkhtmltopdf.
The PDF spec is insane! But like all things what you get out of Word/OpenOffice is 100x more complex than if you wrote it yourself, which is indeed doable.
...a wondeful novel along these lines. They only get shot and kidnapped though. Nothing so bad as PDFs.
I will be extremely surprised if anyone (besides Adobe) has implemented 100% of it.
As far as I know, any PDF can be losslessly converted to an equivalent PDF that can be edited in any text editor, even Notepad. And yes, you could fill in the forms there, too (if you were stubborn enough)
Again, as far as I know, there are no heuristics good enough to get that right for all values of PDF.
The default MacOS PDF printer will actually remap the font cmap making born-digital PDFs where the "text" is something else entirely (say "$" maps to "a").
What? Why!? I've heard of doing that as a form of DRM, but I can't imagine Darwin defaulting to doing that.
This wouldn't be an issue if it was a conscious choice, but when I parsed a lot of born-digital PDFs we ended up with a lot that were like that from various source. Try explaining that...
See http://blog.idrsolutions.com/?s=%22Make+your+own+PDF+file for more info on how to write PDF by hand (also shows why I said you have to be stubborn to do that)
You are trying to move the goal post here. The above statement is simply untrue, PDF is not a text based format and it's that simple.
Look at Chapter 3, Syntax. The code is all text based. We are not talking about the visible characters in a PDF viewer, but the code of the PDF file itself.
BT /F13 12 Tf 288 720 Td (ABC) Tj ET
This can be extended to include spaces so you can essentially mark up entire lines of text at one time. What it can't do is cohesive paragraphs and flow/wrap, you need to use the relative positions to work out what text is in one block (and usually I'd defer to something like pdftotext for simple cases).
Laying out individual characters is common though. It's probably due to kerning concerns.
I silently thank every architect that provides searchable PDFs, it makes my job way easier
Look at Chapter 3, Syntax (and the rest of it, really).
I literally have scripts that use bash and sed, or Python, to modify PDFs by editing the text code. Doing it in Notepad is possible but tricky, as there's a table of object byte offsets near the end that it's easy to mess up by inserting a character.
This is a somewhat big update for PDF.js which is kind of cool in that they haven't really been updating it as aggressively as they usually do in the last year or so.
It's a bit frustrating to work with though. The entire concept of rendering a PDF via JS is fascinating but actually using the API has been a huge pain for us.
We've had to fork it internally and work on typescript bindings and other features to get it to work.
They seem to have a silly policy of only allow developers to use a subset of the API not the whole API itself so that it doesn't look like PDF.js (which I don't understand).
A lot of the functionality just isn't available otherwise.
Quite awhile ago when we decided what parts of the API to version, we thought more people would want to use #1. Now that the project is mature we could probably expose some more base the version off of that.
As for the "so that it doesn't look like PDF.js", we don't limit the API because of this. That suggestion (which I don't totally agree with) came from what we saw people doing, where they'd copy the entire viewer, when it'd probably be better to just let the user's browser choose how to show the PDF.
I'm so sorry about being forward but why the hell don't the vim keys (hjkl) smooth scroll? Its so frustrating. Is there an option to set it as so? Using the arrow keys is so cumbersome.
https://github.com/mozilla/pdf.js/blob/83e1bbea6e23db8744420...
https://github.com/mozilla/pdf.js/blob/83e1bbea6e23db8744420...
All to extract images in a routine fashion.
Just the same, I'm still immensely thankful they've published the library as OSS.
It definitely would have been the better performing route, and simpler to implement. As it is, since the app is low traffic for actual processing of the images it doesn't matter too much, thankfully.
This library was helpful: https://github.com/ScientaNL/pdf-extractor
I haven't had time to do it cleanly, but I should contribute back with the types I wrote after cleaning them up...
There were efforts similar to PDF.js to run Flash content using JS but they were never able to tick all those boxes.
PDF is fine to be some binary blob to download just as most other binary blob formats are.
Would you expect to have .exe files being directly interpreted by a browser?
no, i wouldn't. and yet here we are: wasm.
PDFs are generally an actual document, separate from the site they're on. If images and videos weren't a part of the web pages being viewed, I would be quite skeptical of including them in the browser. I mean, there are JS viewers for STL files (https://www.viewstl.com/) - should browsers include a 2D modeling environment?
> Beyond the core HN crowd, almost nobody cares to have a 3rd party application that they have to install to view PDFs in their browser.
See, I have the exact opposite experience; I've had less-technical family complain to me they were annoyed at Firefox because it stopped just opening PDFs in Adobe and forced them into a crippled slow viewer inside itself. Unfortunately, I can't tell which of us is in a bubble.
> Having a lightweight and secure PDF viewer that is also not made by some 3rd party company that could be collecting any amount of data on you is a good thing in general.
That many PDF viewers are awful is an argument for making a better PDF viewer, but not for baking it into a browser.
See, I don't really agree with that because to me, PDFs are a pretty core part of content on the internet that users browse to via their browser. Pretty much every restaurant makes their menu available on their website as a PDF document. Almost all users will interact with PDF documents while browsing the web at some point or the other. Otoh, a tiny fraction will even know what an STL file is, let alone care about opening/viewing one. So that comparison really isn't a fair one.
> That many PDF viewers are awful is an argument for making a better PDF viewer, but not for baking it into a browser.
That's a bit of an odd statement. If anything, it proves exactly why this is a good move from Mozilla. The PDF standard has been around forever, and yet there is a dearth of free, high-quality PDF viewers that aren't bloated or filled with ads or spyware or trying to get you to upgrade to a paid version of their software. So Mozilla has finally taken matters into their own hands and provided a pretty good, light-weight and integrated solution that will do the job for most users. Power users who care can still enable other software via the plugin system as their default PDF viewer. I'm not sure how you can blame Mozilla for addressing a very real deficiency in the state of available software for PDF viewing.
So for sure you already have access to Evince and Preview.app, they already do everything you want, but Windows users don’t really have that luxury! Being able to say to users to just install Firefox if they want to edit PDF is really good IMHO, way better than the current situation.
Then I had to use Windows. Good god, PDFs are horrible here. No matter what I use, every application is horrible in its own unique ways. Nothing can compare to the default software provided for free with Macs. I'd prefer to manage PDFs on my phone than my work computer.
If Mozilla can help people edit PDFs to any extent, they're doing the world a service.
I have it on my work computer and haven't noticed anything I would rate as particularly obnoxious, but I don't use it much.
Edit: Example, which you may not see if you aren't on Windows. https://musteat.org/images/hn/abobe_install.png
I will use Firefox for editable form pdfs but for those that don't have editable forms, I will continue to use Okular/Gimp.
I actually stumbled across the ability to edit forms in Firefox only recently. I was like... What? This is amazing! For some reason the pdf i clicked on opened in Firefox and yeah, surprised.
It is not something you have to everyday or something, but the existing solutions suck massively. You either have to use Adobe, which requires Windows (or Mac, I suppose) and your firstborn or use some massively shady online service. So personally, I love this feature!
(And I also do not think that this will halt all other development at Mozilla like some comments here imply)
https://www.microsoft.com/en-us/p/okular/9n41msq1wnm8?active...
This is changing quickly with the covid pandemic. One of its silver linings is that I can't remember the last time I had to wait in line for one hour only to be told off by some exhausted bureaucrat about my missing grand-parents' birth certificate or whatever the hell they come up with. Take that, bureaucracy!
A bit off topic from the post at hand, but my gripe was the opposite. The relentless pursuit of parity made them indistinguishable giving users no reason to switch (and taking dev time away from distinguishing features). Granted the pursuit of users instead of principles is its own folly that's hard to overcome when money is needed.
* blank pages when trying to load an imgur gallery on v68 (esr).
* image uploading not working right on instagram and various other sites, either producing blank images or ones with weird lines.
* several teleconferencing / video meeting websites just don't work properly, whether it's not detecting hardware properly, etc
I have to keep chromium installed so I can use these sites properly.
https://bugzilla.mozilla.org/enter_bug.cgi?product=Web%20Com... if you're on desktop.
Like PDF support?
Do you see the problem in all PDFs? Maybe there is something unique to the PDF you are searching?
Any alternative would need some very compelling reason to use it instead. Take Microsoft’s XPS which I think is it’s closest rival. It is an open standard based on XML. It’s built into Windows, Office, and many printers support it natively along with major software vendors, but I can’t think of a single time I’ve come across an XPS file online.
Then, you have to worry about market share and acceptance.
Of course, I do wish Sumatra supported filling forms. Then I could uninstall Firefox too! ;-)
For a local web use, I built for myself https://formulairemagique.fr for this very reason
https://chrome.google.com/webstore/detail/pdf-viewer/oemmndc...
[1] https://gitlab.freedesktop.org/poppler/poppler/-/issues?labe...
Does anyone see a trend moving away from the PDF standard in recent years? Tried to look for data on it but found nothing.
1: https://addons.mozilla.org/en-US/firefox/addon/print-to-pdf-...
Chrome's save-as PDF produces actual text. It's the main reason I still have chrome installed.
Something on your system might be interfering with the printing process.
[1] https://gitlab.freedesktop.org/poppler/poppler/-/issues/463
http://blog.pdfshareforms.com/pdf-2-0-release-bid-farewell-x...
https://blog.adobe.com/en/publish/2017/08/08/taking-document...
The PDF was text-based but every time I copied something it added millions of new lines and hyphens and extra text that wasn't shown on the page.
With that and form fill I basically don't need another PDF reader, which is nice.
https://github.com/mozilla/pdf.js/blob/83e1bbea6e23db8744420...
https://github.com/mozilla/pdf.js/blob/83e1bbea6e23db8744420...
And then what? Fax it? Sounds like a missed opportunity to me. It would be nice if you can add a Submit button to have the data posted to the server, just like any other web-based form.
AFAICT PDF.js is just another JavaScript application and thus as sandboxed as any other website.
They can be part of the initial install so that Mozilla can provide the browser as they envision it, but be able to be removed for those who have other ideas of what their browser should consist of.
I don't know how technically feasible that is with their code, but it makes sense to me from a developer standpoint.
In the sense of a “form” just being lines on paper that you can arbitrarily add some text to — no, that’s easy.
Likewise, in the sense of a “form” being some defined input regions that accept your keystrokes and turn them into new text DOM nodes in the PDF itself — easy enough. Though, unlike HTML, there’s no concept of an <input> tag that just has the semantics of accepting keystrokes and turning them into (persisted) input; instead, this all has to be done through scripting [i.e. writing event-handlers, or having some PDF authoring software generate them]; and there are several incompatible scripting languages for PDF that get used, some of which are proprietary with no open specification.
But, doing form validation? Or, worse yet, making one of those fancy PDF forms that auto-calculates fields like an Excel spreadsheet? Now you’re getting into the hairy stuff, because IIRC none of the open-standard PDF scripting systems provide these sorts of mechanisms, so these are inherently proprietary things.
And when I say “proprietary”, I mean “like old versions of Word or Photoshop, where each version emitted its own in-memory data-structures to disk without formal serialization; and it was the job of authors of future versions to write importers to deserialize whatever format resulted.”
If we simply had print-to-HTML functionality which resulted in a document identical to what you view onscreen while editing, PDF could die the death it deserves.
But HTML+CSS somehow manages to suck just as much for common usage, so it persists.
And implementing a PDF viewer is already a major undertaking; adding the form functionality complicates things even more.
The vast majority of form is indeed not « ready » for input, requiring users to go through hoops to fill them. And that work is done again by the next person.
What’s the use case? Printing out filled-in forms? But otherwise, who would want the PDF in electronic format? It doesn’t seem like a practical way for users to submit data.
Are there utilities that extract PDF field data and submit it to a database? I'd be grateful to see examples.
What about field validation? The PDF may have some minor validation, but that's no substitute for the validation done in a DBMS.
If you want users to be able to save a nice looking form, you'd still want the data entered online directly into a DBMS. I'd offer a "download PDF of your input" as an option, for example.
That said, plenty of users of PDFs have a very paper-based/manual workflow still, and not the motivation and expertise to run and update an online form thing. Or they need to have the ability to handle odd inputs anyways, because paper forms have even worse input validation.
And from a browser/user perspective, the feature here is useful because people expect me to handle PDFs and do not provide nice web forms. They might have terrible reasons for doing so, but I still need to live with that.
A lot of people don't care, because they come from forms in cartaceous - where they have to manually retype everything anyway. For many, their "DBMS" will be an Excel sheet with a dozen rows. The more advanced types likely have some Adobe software that does all the magic.
Fillable PDF forms are really seen as a courtesy to users more than anything particularly useful to the emitter.
Is there a version of Firefox that removes this bloat?
Given that Mozilla is very resource constrained, why are they working on features that aren't necessary?
I think Mozilla's line of thought here is that PDF documents are widespread in the web, to the point where they are a de facto web document type. So it makes sense for a web browser to support them rather than calling out to a user's desktop program (though I assume you can configure it to do so instead).
There's probably a bit of "our competitors do it, so we have to too" in there as well.
And I also like that I can preview long-form PDFs in the browser, before choosing whether to save them and read them “for real.”
Imagine if every time you opened a direct-linked JPEG image in your browser, it treated it as an attachment, downloading it and opening it in your external image-previewer app, rather than rendering it as a synthesized HTML DOM wrapper around the image. Wouldn’t you be annoyed by how cluttered your Downloads directory would get with random files you never actually wanted to save?
And the reasons for not requiring an outside PDF reader are major: It's yet another likely-to-have-vulnerabilities program people need to install, then update. In most cases, avoiding Adobe programs on your PC is a good way to avoid a lot of vulnerabilities.
Chrome/Chromium uses PDFium, not PDF.js, so no. Not sure about Edge.
PDFium has been able to fill out forms for a long time. What’s new for Chrome is the ability to save edited PDF (as fillable).
Having a sandboxed PDF viewer that works 95% of the time is great. For those 5% circumstances where I am actively trying to view a PDF and it won't work in browser, I'll gladly go through the minimal effort to open it in an external viewer.
I don't want to start explaining to my mother, over the phone, how to install and use the pdf viewer anymore :|
PDF is a dork. It's an accessibility nightmare with no obvious advantage over simple ordinary webpages. Somewhere in the comments below, it is mentioned that supporting PDFs is a non-trivial piece of technology. May be! Even steam engines have non-trivial technology under the hood.
It is easy to criticize something when you don't look back at the historical context through which it emerged. It has plenty of advantages over HTML but they're easy to dismiss if you don't have a use case for them.
Can you discuss some of the advantages? The only advantage that comes to mind is that Apple has built-in support for writing PDFs and that has a lot to do with Adobe rather than PDF being a better candidate.
If you think MSOffice users would prefer to output HTML over PDF, you don't live in the same corporate world I inhabit.
I doubt "all" their money goes towards the pdf-reader bit. And tbh, I'd say nobody will really lower their goodwill towards Mozilla because they add features that a lot of people actually need.
When a PDF that has interactive form fields, calculated auto-populated fields, fields that are enabled/disabled according to the inputs of other fields, etc. — the organization that created it (usually government or education) usually does that because they want you to fill it out using a PDF viewer; save it (which will persist the form inputs “into” the resulting PDF); and then submit the modified PDF file back to them. They want this, because they can use automated backend processes to extract the data from the PDF. They don’t want you to just print out the thing and fill it out. In fact, many such “fillable” PDFs start off in a state with many of their form-fields disabled and voided, such that printing them out in that state would result in a form you can’t really write on!
So, at that point, why didn’t they just make the PDF a web page? They’ve essentially reinvented a web form, but with extra steps. The only benefit a client gets is the ability to edit and save the form offline (but that can be done in a browser, too, with local storage); and furthermore, the ability to treat the resulting filled form as a file, moving it around before you submit it. But the cases where you need that are very niche, compared to the cases where you can just direct employees to your Intranet portal.
2. IT person + webserver costs have to included in the budget somewhere. Which can be a big problem.
3. The webpage form can fail, and the support for it has to be provided by the IT dept. If the PDF form fails, dept can handle it on its own, and will often accept a filled+scanned print out of the PDF form.
4. Adding to the point above, PDF forms degrade gracefully, If they don't work, or internet doesn't work, or someone is on holiday, you can still print, fill and hand them in person. Webpages can degrade catastrophically where you whole dept grinds to halt while the IT person tries to fix the problem.