Horrifying PDF Experiments
github.com
github.com
• Just showing the user text? Compiles to plaintext.
• Get the user to give some input? Compiles to a styled form, as PostScript.
• Add radio buttons? Compiles to a physical form but with a 3D-printed notched slider glued to it.
• Require validation for freeform-text form fields? Compiles to a 3D-print + VLSI + pick-and-place specification for a tablet embedded-device that displays the form and does the validation.
Now imagine a "printer" that takes such abstract documents as input, and can print any of these... :)
it doesn't actually execute code, right? Then what's the power of having a compiler in a PDF? you can output the executable, but can you run it?
also, is the "input" and "output" of this compiler just code and executables?
> also, is the "input" and "output" of this compiler just code and executables?
Mostly yes. I'm not sure how much of a typical build chain he was trying to convert to JS here, but the compiler itself typically takes a bunch of files with C code and outputs a number of "object files", which are really chunks of machine code. In an actual build process, you'd then use a linker to glue those object files together in the right way and make an executable.
I guess, what you could do if you wanted was to include the whole build chain (including linker) into the PDF, encode the executable as Base64 and fill some form field of the PDF with it. Then your workflow would be as follows:
1) Write some C code
2) Copy the C code into form field #1 if the PDF.
3) Hit a "compile" button or something. Form field #2 fills with what looks like an enormous amount of random gibberish (really the Base64-encoded binary)
5) Copy all the text from form field #2, use a program of your choice to decode the Base64 and save the decoded binary on your hard drive.
6) Run the binary on your hard drive and watch your C code execute. Hooray!
Depends what you mean by "run", really. You can write a full-on X86 emulator, and execute a compiled binary there. But given that it's an emulator running in a nested series of sandboxes, it won't be terribly useful -- for example, it still won't have I/O capabilities.
Does this all sound like a fantasy? It should, because it is. Absent the history of it actually happening and being able to point at that, the question is akin to comparing two sports teams across history, eg the 2014 Golden State Warriors to the 2002 Mavericks and trying to talk through which team would win.
Could an independent Macromedia have been better stewards of Flash than Adobe, leading to a world today where Flash wasn't deprecated? Absolutely. Would it have? We'll never know. Flash had a number of issues that lead to its death today, and it's not clear if an independent Macromedia, with a different internal developer and business culture from Adobe, could have fixed all of them resulting in a different future, or if they even needed to be fixed for that future to happen.
Looking at Adobe's poor stewardship of PDFs, however, it's hard to see positives to Adobe-owned Macromedia and Flash.
Flash could have become an open web standard
driven by a programming language that isn't
javascript, which we're all now forced to use
due to browser support.
Agreed on the impossibility of discussing what-ifs.Obviously, they could have done anything. =)
Ultimately though I guess what I'm ultimately asking is if there were any hints about how Macromedia would have done things differently, had they remained independent, particularly in the direction of making Flash an open web standard.
The charts were actually executable Postscript, running the algorithm.
One of the coolest things I ever saw.
Imagine digital academic "papers" in STEM fields that natively ran the simulations the paper was describing. Jupyter sort of delivers that, but it still feels like early days for interactive digital-first documents (or as Steve Jobs has been credited for saying, "bicycles of the mind").
And sure, some simulations are very heavy, but they are more of exceptions. Also possible to have the best of both worlds, and have both a simulation, and a static snapshot available.
Ability to rerun programs is great, but we should be careful to remember that it's a different thing than reproducibility.
Imagine trying to figure out some 2001 JS paper thing for ex. But applied to every generation of technical development.
There’s always standards of course but we’ve seen those go sideways enough time to make one cringe at the thought of ‘dynamic papers’ via some new medium.
The kind of thing that sounds amazing on the surface then you remember the sort of crazy IT depts that thousands of universities run and forget the whole thing.
I suppose this only works in a few fields though.
It shouldn't be! Reproduction needs to involve the interaction of human brain meats with a human level description of the solution. This is how we make sure that people aren't talking about something different than what was actually done, and how we make sure our conclusions are robust against the things we've failed to specify.
Imagine saying the same thing for physics: I start replication by running a time machine and using the same apparatus as the original experiment under the same conditions. Impracticality aside, this would be potentially useful to suss out fraud and certain kinds of errors, but what successful replication tells is is manifestly less powerful than successful replication on a new apparatus in a new location at a new time, with new values for everything we've failed to control.
For things that are heavily resource constrained, it still could be a boon to have interactive access to the data that comes out of it.
But-- wouldn't it be cool if there was a way ordinary people could create interactive content to interact with data in a rich, intuitive way?
1. On the creation side, it requires someone be comfortable with Python (or other Jupyter language) to some degree. Right now, programming is still considered a career skill rather than something "ordinary people" should be expected to know. Perhaps layering a graphical programming interface on top of this, which UE4 seems to have had some success with with their Blueprint system, would get "ordinary people" over the mental hurdle of being intimidated by code-as-text. Just look at the mental gymnastics people will engage with in Excel while thinking it's not programming.
I see this as more of a social problem than a technical one, at any rate.
2. Once you build an interactive Jupyter document (especially if you use interactive widgets), it's not necessarily that easy to share in its original state without requiring the reader also have a Jupyter environment set up or access a server running Jupyter. I would like to be able to share the document in a way that can be accessed offline by someone without them needing to set up the whole environment. Maybe an "Adobe Reader"-like application for Jupyter notebooks that "ordinary people" can just install with a click?
#2-- Or just use the browser. It's capable enough, even if large datasets are somewhat problematic. The hard thing is the UI and identifying what the correct subset of functionality to surface is.
Because it's not installed, and they don't want to and shouldn't have to learn something new when there's something not new already at hand which suffices.
If you ever find yourself saying something like, "people can just do X" and wondering why they don't, turn it around and ask yourself, "why can't I just do Y?" In this case, that would be, "Why can't I just make my notebooks work in the viewers that everyone already has agreed upon using (i.e. the WHATWG/W3C hypertext system, i.e. the web browser) instead of asking them to futz around with installing and learning Jupyter?". When you start making excuses for why not, it's the moment you should be able understand another person's reasons for why not Jupyter.
Reason it's done in pdf is a lot of our technical is spat out in PDF format (generated from CAD - SolidWorks).
There are other options like Traceparts or setting up a variable input SolidWorks model to generate loads of static outputs, if you have the time and money.
- GIF is obsolete (~100x heavier than MP4 in my use-case, so out of the question)
- MP4 has poor support in PDF readers
(- Besides, PDF is not appropriate for electronic documents.)
- EPUB doesn't seem to support MP4 at all
(- EPUB does support PNG, not sure about APNG, will have to try it out...)
- MHTML=EML support has been dropped from browsers, which is completely baffling to me. There are alternatives like SingleFile, but they feel like dirty hacks : https://addons.mozilla.org/en-US/firefox/addon/single-file/
- What future for AV1 support ?
What I've read of EPUB is also pretty disappointing. Seeing as it's a compiled format, once again, instead of going the zipfile + bunch of html inside + specific layout, we could have had a subset of html in .mhtml.gz with, like, metadata in a <script type="application/json" id="x-epub-metadata">. And then, guess what, web browsers could have been able to read it natively…
Probably just some special code in Windows Explorer watching for the combo of .htm(l) file plus simlarly-named folder – via the command line I can delete just the HTML file or the folder separately without problems.
Yeah, if I'm not mistaken, this is what SingleFile uses ?
Well, Libre Office Writer deals with (multiple, 100 Ko < size < 10 Mo) MP4 just fine. It's when the ODT is converted to PDF that most(?) PDF readers seem to be unable to read those MP4 properly.
1. using vector graphics wherever possible and then encoding it as SVG
2. if bitmap graphics are absolutely required and they can be procedurally generated, then do that
3. if large photographic data, video, or any other kind of data is required that can't be handled using the above steps, then separate that data set as you normally would using the file system directories, place the data set subtree into a ZIP archive, write your code so it references items by file paths relative to the ZIP, and then put your page into the root of the ZIP file, too, e.g. as index.html—your readers and reviewers follow along by using their system's native ZIP support to explore the contents of the ZIP file so they can locate index.html and then double click it, and index.html opens up with an "open dataset" button which you use to then feed in its own parent ZIP archive
The last part might sound complicated, but it's not much different from asking someone to use MS Office or VS Code or an IDE to open a file/project. (It's just that instead of requiring then to already have that IDE installed, you're giving them the IDE they need at the same time that they're getting the document/dataset they're actually interested in).
These approaches are robust enough that they're very unlikely to be broken by future browser changes. It's not that the tech is lacking right now, it's that human habits are lagging behind and we haven't yet established this as a cultural norm/protocol/expectation.
†IMHO as long as your document doesn't cross 10 Mo, you shouldn't have to separate the data…
The packaging convention I described is similar to the container formats used and created by MS Office apps. The difference is that DOCX, XLSX, etc rely on XML instead of HTML that can be used without requiring a separate proprietary app. People create and exchange those files every day (even for things as trivial as a single-page flyer) without knowing or caring about whether it should "warrant this kind of treatment". Worrying about a purported edge case for <10 MB(?) of data sounds like an imaginary concern.
Getting this to work with Latex was... interesting. I spent a lot of time typesetting as a grad student.
[1]: https://www.sumatrapdfreader.org/free-pdf-reader.html [2]: http://www.finalmesh.com/ [3]: https://khoadabest.surge.sh
mutool clean -d your.pdf clean.pdf
Now open clean.pdf with a text editor.
Thank you!
I'm personally idly curious, but have no experience with reverse engineering or 3D or file formats... so the emphasis on my end is idle curiosity :). But it's possible that many such people poking around may still generate interesting leads.
Depending on how effectively intraoral scanners can scan things other than teeth, offering to scan random objects people send/bring in, on a best-effort/no-warranty basis, may also generate practical interest.
(Also, wow, looks like these things are in the $25k range?)
Also, a very small extra thought, scanning extremely simple objects like cubes and flat planes may make the analysis process slightly easier because the data in the file will be easier to pars--wait. Okay I have more ideas.
Can you convert/import arbitrary 3D data into the proprietary dxd format? If there is any way to do this, there is nothing else that will move the analysis process as far forward as quickly, and offer the best chances of producing the most complete result. This is because a) the data files will have 99% less complexity due to being synthetic and not containing noise associated with real-world data, b) they'll be full of reference points from known 3d models, and c) entirely controllable input data gives the highest chance of figuring out all values/fields in the model files.
If this is possible, chances are most imports would be user requests based on the analysis process ("does changing this value alter this byte?"). Initial ideas I can think of would be the 3D Teapot, a single pixel :D, and simple cubes, triangles and planes.
Lastly, coordinating a backup installation of the scanning software onto a dedicated machine, or moving the main install onto such a machine, that enterprising reverse-engineers could connect to remotely (ideally at any time of day, and obviously after privately negotiating credentials) and install debugging tools (read: IDA/Ghidra/etc) onto, would likely be extremely helpful, and should provide the best "how we reversed this" narrative with regards to licensing. This would simplify the import request situation too.
If importing is not possible, IDA et al may end up being necessary to understand certain complex details or possibly even get started. Solving the "generate interest" problem would naturally be more complex in this scenario though. :/
I think I've really exhausted my knowledge in this area now :), although I do remain interested in knowing how things go.
I published all data in this Github repo: https://github.com/thangngoc89/dxd-file-format
I also tried to scan something simple like an sphere or a pencil without any success. The software only recognize tooth-like structures.
Luciky, it can exports to STL files with 100% triangles that can be imported to others dental CAD software so I hope it would help with the progress.
It's regrettable but understandable that the software only recognizes/accepts teeth considering the postprocessing it clearly does.
And CC0ing the model data is pretty much the textbook approach to analytical freedom :)
(And just to confirm, STL/etc->DXD isn't possible?)
Thank you for reminding me about this.
Quick update: opening the DXD file with a hex editor, there is a XML file defining the metadata of the current file and a public RSA 1024 key. I’ve been scouring around to find the private key with no success.
Hmmmm. Ideally that key is only being used for attestation/authentication, not encryption. In this case, you definitely don't want to locate the private key, because that key's confidentiality is what verifies the integrity of the scans made by your device.
Also, said private key might be specific to your copy of the software to create a chain of custody to your machine for medical purposes, or even more likely for licensing reasons.
In any case, if it's being used for encryption, that would amount to an unfortunate DRM situation that might be a bit of a hornet's nest to fiddle with, because of the high likeliness the key is being used for license enforcement etc (tracking scans made by copies of software deemed illegitimate etc).
It's very cool you can go from STL to DXD though. Now I'm curious, was the STL file that crashed the software originally generated from a DXD file created by the software? It originally being a DXD should be irrelevant, but chances are the pipeline inside the software chokes on things that aren't models of teeth. This does admittedly make the reverse engineering process trickier...
I am sure there are libraires un python or JS nowadays, it’s just a question of parsing the tree to find the u3d node and dumping it out, very simple
Edit/Appendum: Crafting an interactive website (i.e. without the dependency on jupyter) might be more future proof.
I personally prefer Jupyter because it seperates the programming language (Python, R, C++, etc.) from the representation (for instance Web) and still allows interactivity for a certain degree (given a backend running the source code).
Putting your animations in traditional video formats (mp4, ogg) or on vimeo/youtube is probably the best way to make them accessible for most people. Many scientific labs have their own youtube channels.
(And a website doesn't fit the requirements as it's not contained in a single file, so its archival is a lot more complicated.)
Back in 2011 I used it to make a whole lot of figures for a multivariable calculus course; they're still in use.
Want to include the CSV raw data with your report? Just add it as a PDF attachment.
Want to hide a game with your homework? Add it as a PDF attachment. Chrome and Preview on Mac doesn't show that it exists, but Firefox can be used to extract the file.
It's not going to shock anyone to have a 5 MB file as a PDF, but there's a lot you can hide in there (MP3s, games, HTML files including more JavaScript, whatever else).
On the surface, everyone thinks it's just another PDF. But the real data is hiding in plain sight.
It's probably underutilized overall, but there's nothing hidden about it when most viewers show the data.
Do "Most viewers" show it? Google Chrome doesn't, nor does Preview on a Mac. There is no easy way to add attachments to a PDF, except Adobe Acrobat Pro or iTextPDF. Firefox and Adobe Reader can read the attachments, but it's "hidden" to some degree, inside slide-out side menus. Certainly enough to avoid a casual glance.
If I see a PDF containing a page or so of text and it turns out to be several MB, I would become a little suspicious. But you're right that most people are not aware of the general size of things.
and it gets replaced with a basic filled and bordered rectangle.
...also known as a pixel ;-)
PDFs are a great format if you just ignore the dumb parts. :)
Adobe does have tools for creating PDFs that are accessibility-friendly, but it can take hours of work. As much as it sucks for certain audiences, it just doesn't make sense to do that in the general case.
Hypothetically a compliant reader is supposed to ignore non-PDF/A features encountered in files that declare themselves as PDF/A, so I've sometimes wondered if a cheap form of "sanitizing" PDFs would be to simply force their PDF/A flags on.
That is one shitty site. Trying to shove Google Analytics down my throat, no contact information, no privacy page. Probably illegal under GDPR.
> so I've sometimes wondered if a cheap form of "sanitizing" PDFs would be to simply force their PDF/A flags on.
That's not really how PDF-standards work. You'll have to "rewrite" the problematic parts, the standards are just for checking against the pre-defined ruleset.
In professional media production we do this "rewrite" all the time (PDF/X-standard). Though sometimes PDF files are just so "broken" that it's impossible to fix them.
Yes, I don't think it gets much attention - I should probably have pointed at the github org which is reasonably active. https://github.com/verapdf
> That's not really how PDF-standards work.
Well, it is how the standard works (don't make me dig out the relevant bit of what's publicly available from the standard) - the issue is whether common PDF readers actually do what they're "supposed to" or whether they just try and interpret as much as they can.
As a newbie developer I decided to use PostScript to generate badges for all our employees. There was a list of employee names in a text file, there was a PostScript file with the program and a Perl script to join them together.
The PostScript program would take the names, generate 8 badges per A4 page, scale the name of the employee so that it fits the space perfectly, generate procedural background, etc.
This is like flash player all over again. No way am I going to enable the proper pdf reader for web content view. There's a good reason browsers refuse to support all this
The underlying JavaScript-like language is ActionScript which was originally developed by MacroMedia to provide animation for Flash.
It's quite useful for creating PDF SmartForms that adjust their contents based on the user's responses. Until very recently they were only viable in the official Abobe Reader until Chrome decided to add support.
As far as security, blame Chrome for not incorporating an opt-in before allowing a particular PDF to run ActionScript in the browser.
It is just a scripting language. So did they actually use flash player tech?
[0] https://github.com/QubesOS/qubes-app-linux-pdf-converter
[1] https://blog.invisiblethings.org/2013/02/21/converting-untru...
From the last 7 weeks commit history, I see only people associated with chromium. So may be foxit is not involved in developing pdfium anymore, or may be they're not developing in open and sync in once in a while.
doesnt a compiler just output executables? would we be able to run these executables? where would these executables get stored?
So if you do such a thing, realize you might only be doing it for your own enjoyment.
There’s always a scary world lurking underneath it seems.
Wasn't there a search engine built into finding redacted PDF content? I think it made the headlines here a while back.
"PDF drive" (https://news.ycombinator.com/item?id=25240373, 0 comments) just appears to be an ebook crawler over in the less-than-#FFFFFF-department if you get what I mean.
I also found a thread talking about searching PDFs for specific queries (https://news.ycombinator.com/item?id=10154527) which appears to have generated some interesting results back when the thread was posted, in 2015.
Not seeing anything recent though. But on the subject of a search engine specifically for finding redacted content, I couldn't help but imagine the discussion...
"Hi, I would like to find a •••••••."
"You specifically want a •••••••?"
"Yes, literally."
[Person 2 walks away scratching their head wondering what person 1 would do with a 'hunter2']
However, if you are using a tool like a redaction tool then the software should forbid you from writing in append mode. This was a common error in old PDF apps and perhaps contemporary ones that are new.
Edit for politeness:
My surprise is aimed at the researcher, not you :)
They were just sharing something they were surprised by/interested in. I was surprised to read that's how editing PDFs works too.
I do agree that it's surprising behaviour to regular users of PDF that it usually maintains a history of sorts. Apps should make this clearer.
https://www.merriam-webster.com/dictionary/you:
1. the one or ones being addressed
2. ONE sense 2a (which is “being one in particular”)
So, pro tip: in chat-like discussions with strangers such as hacker news, one should prefer saying “one” when using sense 2, even if it sounds a bit archaic (at least to me. Is it?)Also, when reading a “you” that could be interpreted both ways, do not assume it is used in sense 1.
Aside from being an excellent pdf manipulation library it also has a mode where it outputs a version of the pdf that is much easier to manipulate with a text editor and then lets you build a new pdf from that.
Shout out to Jay who has been steadily working on it for many many years. He is the most kind, undestanding and hard working free software developer I've had the pleasure to cross paths with. Thanks for all your hard work Jay!
QPDF is not a PDF content creation library, a PDF viewer, or a program capable of converting PDF into other formats.
That doesn’t appear to be true exactly but isn’t anywhere in the realm of impossible which is a serious issue for PDF.
I was hoping FoxIt dropped a lot of the BS spec parts, but it seems they don’t want to “lose out” to Acrobat in the features checklist. At least I know it’s easy to turn JS off by GPO with FoxIt, Acrobat I assume too?