Show HN: Paper discussing JSON-compatible binary serialization formats
arxiv.org
arxiv.org
This was very timely for me as I've been actively researching and applying json and binary serialization for the last 2 years, as part of my day job, and also due to a personal interest.
I ended up not using binary serialization/deserialization, so all the following is somewhat anecdotal to the paper itself.
But I'm very keenly interested in those subjects anyway, since at work I also have to deal almost every week with binary serialization issues in at least 3 other contexts and it's a PITA, something which I'd like to solve.
In recent years I've been using json files at work as a means to supplement our internal database-based app at work.
The database is kept in a relational backend, but is deeply hierarchical. It might have been better served by migrating it to a graph-based database engine. Unfortunately this is a major undertaking and and I'm not even sure it would be advantageous.
Currently we've added json to the application in 2 ways:
1) As a way to add version control to our (central, remote) SQL database, by transforming it to and from a repo consisting of json files, where each json file contains a root entity and its hierarchy. I described that briefly here [0].
The main motivation here was using a readable representation of the data, to allow users to see changes in the repo history and merge possible conflicts.
2) As an optimization: to speed up a certain operation that is done repeatedly on this data, by reading some of the data from those json files instead of reading it from the database.
For both options, other formats were considered: using XML, binary or a home-grown serialization.
In fact use-case (1) was previously implemented using a home-grown serialization format, but that was scrapped, partially due to time constraints, and partially due to the fact a whole UI had to built just so that users could look at the repo history, and this made it very expensive and bug-prone.
However I ended up choosing json due to the fact it's a popular AND readable storage format, and I'm counting on future optimizations for parsing it.
For use-case (2) a binary representation would have made much more sense. I had a hunch it wouldn't matter much, so a member of my team implemented an optimized serialization/deserialization implementation of some part of the data in C++, while I used RapidJson. I didn't even use SimdJson due to the fact I still wanted to support older machines, I used the older RapidJson.
We measured the timings, and indeed the binary deserialization was faster, but not by a big margin. And compared to reading from the (central, remote) SQL database the difference between using binary and json was negligible. In the end we decided to go with the json deserialization partially due to the schema upgrade problems and partially because we thought that if the performance difference was negligible then a more readable representation is better.
Although, I didn't actually read the whole thing with RapidJson, I used its "streaming" capability and wrote my own small state machine, so I could skip decoding some parts of the json file I was not interested in.