Sure we do. There are plenty of proofs out there that pi is an irrational number.
It's then trivial to see that every number you can think of is encoded in there, and therefore any data, piece of music or movie that ever existed.
(I'm not sure we're allowed to fiddle with the encoding, but since we allow ourselves to represent a piece of music into a number, we're already talking about encoding anyway, so it doesn't seem like cheating to me...)
Also, 1.01001000100001... is a good example of a number that is both irrational and transcendental but not normal.
According to wikipédia, you gave a definition for "simply normal", and for normal numbers the distribution of any sequence of digits is uniform. So 00, 01, ..., 99 each occur uniformally too.
We don't know whether pi is that either for any integer base.
The first place a valid PDF could be ended, perhaps.
you're just pasting random python snippits at me now. It's time to move on.
again, just to summarize: PDF files do not have to be zero aligned, and they do not have to be end aligned. Therefore the answer to the question "what is the first segment of Pi that is a valid PDF file" is trivially (0,infinity). That is a correct statement. The non-greedy (in the regex sense) answer to that question will be different, however.
Note the word "next", implying that (0,10) sorts before (0,11); you even say it yourself "11 is less than infinity". Where I'm from "first" and "less" are related (the first element in a unique sorted list is defined to be less than all other elements). So if there is any valid pdf in pi that can be identified by the range tuple (0,N), then the first valid pdf must occur before N -> infinity. Therefore (0,infinity) can never be the first valid pdf, even though it may be a valid pdf.
Maybe a picture would help:
Potential pdf file ranges in pi: [(0,0),(0,1),(0,2),(0,3),(0,4),...,(0,N-1),(0,N),(0,N+1),(0,N+2),...,(0,infinity)]
Is it a valid pdf? no no no no no (no) no yes yes yes (yes) yes
Which one is first? ^^^
I thought linking to a python script that shows the order comparison of a tuple (0,N) as less than the tuple (0,N+1) would clearly demonstrate this, but it appears to have failed to communicate that to you. We don't need non-greedy regex rules to do a less than comparison.Including a PDF that generates the digits of pi
I understand the point that PI contains every possible piece of information, theoretically.
However, the chance of finding a given string in PI depends on the string’s length. The longer the string, the more the probability tends to 0.
The paradox therefore is that PI contains every PDF, but you will never find them, so in what sense does it really contain them at all?
Are you saying that:
- given a long string, we might ask “can this string be found in PI?”
- the probability of finding a long string in PI is infinitely small
- the number of possible strings in PI is infinitely large
- it’s not possible to decide if the answer is yes or no?
Probably won't send it to any recruiters, but it will be a funny anecdote for interviews.
And even places with strong roaming rights tend place limits on well marked land.
[1] https://www.pdflib.com/pdf-knowledge-base/pdf-password-secur...
I've read a few comments on HN how PDF is, well, not developer-friendly. If people are interested in providing some more examples here, I'd be curious to know!
The old .Doc and .xls files were a bad format, but my understanding is that since Office 2007 the format is generally much better.
Streams just suddenly end? That’s ok. Totally corrupt xref tables? Ok. Incorrect image headers? Ok. Unrecognisably mangled Type1 font formats? Fine!
I personally think the structural PDF format is a really great format. It's entirely ASCII-based, a pure text format, yet it can embed arbitrary binary data and compress that data. The actual structure is simple and support just enough functionality, like a tree of object, dictionaries and arrays, unicode strings, date formats, etc.
I think if you limit yourself to pure structural PDF woulde have been a great format to standardize upon, much better than JSON or XML. It';s richer than JSON, simpler and saner than XML. Again, it's top-notch ability to embed binary is great. It has other great characteristics, for example you can update anything just by appending.
The ugly bits are in the "semantic" PDF: the page descriptions, media, etc. Even then, the early version of PDF were nice, mainly just simplified Postscript.
A common case is clients who use utilities that generate single customer documents then merge them into a bigger file for bulk print and mail (bills and statements, not identical copy). Without fail that results in thousands of similar but different subset fonts whereupon most printers I've encountered eventually fail due to memory issues.
Typically this leads to a discussion about "I can open it on my computer fine" and bending over backwards to find a workaround. Merging and consolidating these fonts doesn't seem to be a simple task, although some tools claim to work some of the time.
Something that scopes object resources for disposal could be nice in the PDF spec (maybe it exists), but something like a LRU caching mechanism on the printer would potentially resolve this too.
- Clipping path logic, you can write text outside of it, which makes it effectively invisible yet it will show up if you try to extract the text.
- Anything regarding the graphicstate stack, it's a pain to debug.
- Extracting content from AcroForm/JS "XFA" forms
PDF is great format for printing, it's just a pain for pretty much everything else.
My other one is the use of multiple subset fonts that are actually the same font with a different subset of glyphs that you want to merge back together.
Having said that (and worked on a commercial PDF library), despite all the cruft that came with age, it's a well built format that survived the test of time with good reasons.
I worked on a Pdf with inbuilt tracking solution that updated the form layout using ActionScript based on the workflow status and the role of the user (ie the line manager had a different group of fields in the form to complete like their signature while viewing what the initial requestor had entered) Lots of callbacks to the server saving in progress data and updating the status of who had the form, who was next based on department and emailing it to that person if it passed validation.
An initial fun discovery was that you could force the form to download and replace itself with the latest version even if they had just opened some old file they had on their pc.
EDIT: oh yeah, I'm pretty sure it contains Mail, just like Zawinski said
A lot of the time developers want access to the text inside PDFs. Unlike HTML or formats like MS Word (XML or old binary format) getting "text" isn't really possible.
Most "document" formats have the concept of words or strings: a set of characters separated by whitespace. PDF isn't a "document" format in that sense - it's a page description language. Instead of strings of text you have character glyphs positioned at a particular location.
If you want to "read" the text, you have to work out the orientation (which can change throughout the page - think of table header alignment), and use some kind of heuristic to guess the word spacing based on the font and character spacing.
There's also this whole thing with clipping, where some text can be hidden behind other objects (or off page) so you have to try to deal with that.
There's lots of libraries that try to do this for you, but there are lots because none get it right 100% of the time...
Well, at least, pdf is probably better than printed paper for that purpose.
That is truly nightmare material. Especially considering a non-trivial percentage of pdfs circulating are non-conformant and people still expect them to render...
Mozilla's pdfjs[1] project is a pure HTML/JavaScript solution for PDF rendering. This is the same code that ships in Firefox browser as well. This is standalone, AFAIK, it doesn't talk to a mothership.
Ended up just making the app generate HTML before calling wkhtmltopdf.
The PDF spec is insane! But like all things what you get out of Word/OpenOffice is 100x more complex than if you wrote it yourself, which is indeed doable.
...a wondeful novel along these lines. They only get shot and kidnapped though. Nothing so bad as PDFs.
I will be extremely surprised if anyone (besides Adobe) has implemented 100% of it.
As far as I know, any PDF can be losslessly converted to an equivalent PDF that can be edited in any text editor, even Notepad. And yes, you could fill in the forms there, too (if you were stubborn enough)
Again, as far as I know, there are no heuristics good enough to get that right for all values of PDF.
The default MacOS PDF printer will actually remap the font cmap making born-digital PDFs where the "text" is something else entirely (say "$" maps to "a").
What? Why!? I've heard of doing that as a form of DRM, but I can't imagine Darwin defaulting to doing that.
This wouldn't be an issue if it was a conscious choice, but when I parsed a lot of born-digital PDFs we ended up with a lot that were like that from various source. Try explaining that...
See http://blog.idrsolutions.com/?s=%22Make+your+own+PDF+file for more info on how to write PDF by hand (also shows why I said you have to be stubborn to do that)
You are trying to move the goal post here. The above statement is simply untrue, PDF is not a text based format and it's that simple.
Look at Chapter 3, Syntax. The code is all text based. We are not talking about the visible characters in a PDF viewer, but the code of the PDF file itself.
BT /F13 12 Tf 288 720 Td (ABC) Tj ET
This can be extended to include spaces so you can essentially mark up entire lines of text at one time. What it can't do is cohesive paragraphs and flow/wrap, you need to use the relative positions to work out what text is in one block (and usually I'd defer to something like pdftotext for simple cases).
Laying out individual characters is common though. It's probably due to kerning concerns.
I silently thank every architect that provides searchable PDFs, it makes my job way easier
Look at Chapter 3, Syntax (and the rest of it, really).
I literally have scripts that use bash and sed, or Python, to modify PDFs by editing the text code. Doing it in Notepad is possible but tricky, as there's a table of object byte offsets near the end that it's easy to mess up by inserting a character.