Reverse engineering the Notability file format
jvns.ca
jvns.ca
Reminds me of this (it gets even better in the comments): https://thedailywtf.com/articles/XXL-XML
Nothing in here was really complicated – it was just some existing standard formats (zip! apple plist! an array of floats!) combined together in a pretty simple way.
It's rather telling that wrapping points in 3 layers of encodings is considered "pretty simple", almost the norm today, when years ago doing such a thing would probably elicit reactions more towards WTF.
• Likely this file isn’t literally a PList, but the output of a data storage framework (e.g. CoreData) that uses PLists as one possible backend. In the code, it just looks like adding annotations to native data structures to make them “storable” within a database-like object.
• Only the text-formatted PLists contain base64’ed binary data. The binary-formatted PLists just contain length-prefixed binary data. There isn’t the “binary in base64 in XML in binary” that you’re imagining.
This explains the feeling throughout of "this doesn't feel like a format the app developer designed".
I remember trying to reverse engineer older Symbian notes format -- it was a custom datastore format containing various binary "chunks", each with semi-custom compression. I had to recombine multiple chunks just to get a formatted text, and I still had occasional failure.
Or look at other vector format -- like CGM, first made in 1987. It has its own serialization mechanism, it's own model (file->body->figure->segment), and 3 compression types specifically designed for it. Look at http://standards.iso.org/ittf/PubliclyAvailableStandards/c03... -- do you want this today? I'd that zip of plist of float anytime over that.
To get a feel for her work, a bunch of her posts are here: https://jvns.ca/categories/favorite/
In the end I had to compare a working english pack and the disfunctional one; the actual issue was something pretty simple - the names of the properties in their ini file were bad for some languages, and some checksums were wrong, or something like that. But it made for interesting black-box experience; there's of course no feedback about what went wrong in the device, nor any accessible logs :P
Today, many programmers tackle the challenge of ad hoc data by writing scripts in a language
like PERL. Unfortunately, this process is slow, tedious, and unreliable. Error checking
and recovery in these scripts is often minimal or nonexistent because when present,
such error code swamps the main-line computation. The program itself is often unreadable
by anyone other than the original authors (and usually not even them in a month or two)
and consequently cannot stand as documentation for the format. Processing code often
ends up intertwined with parsing code, making it difficult to reuse the parsing code for different
analyses. Hence, in general, software produced in this way is not the high-quality,
reliable, efficient and maintainable code one should demand.
[0] https://www.cs.princeton.edu/~dpw/papers/ddc-jacm.pdf (pdf)After finding a couple hints on what WndProc messages the control would respond to, I spent a few weeks reverse-engineering it, and wrote up my findings here: https://groups.google.com/d/msg/microsoft.public.dotnet.fram... .
I never did wind up building anything useful with that information, but it was a pretty fun exercise.
It's kind of a tangent but it's somewhat related.