People will disabilities that rely on screen-readers don't love it. There is no such a problem with HTML/CSS which should be the norm for internet documents.
> With PDF, everyone sees the same thing.
Yes, provided you can see at first place...
[1] https://developer.mozilla.org/en-US/docs/Web/HTML/Element/Im...
Why would I want document type that can't even refloat on display size to represent any longer written text that is supposed to be consumed on digital device?
And as a format, it's much more sane than, say, Word or Excel.
Even if the focus on "where do I put this glyph" means the original text isn't in there by default.
FWIW this is not a technical barrier; it would be absolutely trivial to associated blocks of non-flowed text with the layed out text.
It should be. But meanwhile everybody seems to think it is perfectly ok that there is a bunch of JavaScript that needs to run before the document will display any text at all and how that text makes it into the document is anybody's guess.
My bashful take is that nobody told the rest of the web development world that they aren't Facebook, and they don't need Facebook like technology. So everyone is serving React apps hosted on AWS microservices filled in by GraphQL requests in order to render you a blog article.
I am being hyperbolic of course, but I was taken completely off guard by how quickly we ditched years of best practices in favour of a few JS UI libraries.
You like it for that very quality: immutable, reproducible rendering.
Those who have to extract data from PDFs face nearly the same problem as those who have to deal with paper scans: no reliable structure in the data, the source of truth is the optical recognition, by human or by machine.
"Because it's pretty" is it. 99% of people don't care about text being a data mining source.
I quite like PDFs, but this thread has been an eye-opener.
I do wish they had focused a bit more on non-visual aspects such as screen-reader data, but to say the whole point is "because it's pretty" is a bit uncharitable. The format doesn't solve the problem you wish it solved, but it does solve a problem other than making things "pretty."
I agree, I do too (LibreOffice), but for the opposite reason. Even internally, the font rendering in LibreOffice with many fonts is often quite bad. This is especially noticeable for text inside graphs in Calc.
If I'm going to read something lengthy that's a LibreOffice document, I open it (in LibreOffice), and export it to a PDF. LibreOffice consistently exports beautiful PDFs (and SVG graphs), which tells me that it "knows" internally how to correctly render fonts, just that its actual renderer is quite bad.
The other thing that is unfair is assholes who deliver tabular data in PDF format usually don’t want you to have it. When your county clerk prints a report, photocopies it 30 times, crumples it and scans to PDF without OCR, that’s not a file format issue.
I sometimes see people complain about how PDF sucks because it doesn't look quite the same everywhere (namely, non-Adobe readers), but if you're not doing anything fancy is pretty much does. It is, at minimum, more reliable than any other "open" format I'm aware of, save actual images.
I know something about that area. Today, perhaps a 10th of CVs are sorted and prescreened by software. That fraction will only increase.
There are two issues with parsing them however.
1) PDF is an output format and was never intended to have the display text be parseable.
2) PDF is PostScript++, which means that is is a programming language.
This means that a PDF is also an input description to the output that we
are all familiar with seeing on a page.
PS I don't know if it is the case anymore, but Macs used to have a display server that handled all screen images in PDF format. That was an optimization from the NeXT display server, which displayed using Display PostScript.Quartz! https://en.wikipedia.org/wiki/Quartz_(graphics_layer)#Use_of...
The big change that came with PDF was removing the programming capabilities. A PDF file is like an unrolled version of the same PostScript file. There is still a residue of PostScript left but in no way can it be described as a programming language.
A programming langauge is not inherently a programming language due to features it contains but due to it being used to program. A program is "a series of coded software instructions to control the operation of a computer or other machine."
In this way, a PDF file embodies a program that performs specific tasks. A PDF file does not contain a general purpose programming language, but it does contain the page description language of the output format that describes what is to be imaged. Then, the PDF program is given to an interpreter that displays the output.
This is the same as a simple program in turtle graphics to display a rectangle, even if no other language feature was used. In such a case, one would say that rectangle was programmed. We would not use the word program in connection with that turtle graphics program, if the rectangle description were not sent to an interpreter that displayed the rectangle.
SVG does that too, but it can also have aria tags to improve accessibility and have text that can be extracted much more easily.
Mostly. I've seen issues where PDF looked fine on a Mac but not on Windows.
Also, the fact that you see the same thing everywhere is good if you have one context of looking at things - e.g. if everyone uses big screen or if everyone prints the document, that's fine. But reading PDFs on e-book readers or smartphones can be a nightmare.
Why should they all see the same thing?
Using PDF here is self serving. It’s actively user-hostile.