JCOF: JSON-like Compact Object Format
github.com
github.com
* JSON is really schemaless so I don't have to assume all objects are shaped the same or that they are even the same kind of object. This allows for streaming serialization, and does not require the data structure to be known or to introspect data to create the heading lines.
* Nested objects look to be difficult, especially if the schema is not known.
* JSON is very human friendly, JCOF is not.
I do think the format is better than CSV, TSV and other delimited flat record formats. JCOF appears to be able to handle objects that are not tabular, which CSV just can't really do. A replacement for JSON this is not, and storage efficiency isn't really relevant because of the difference in purpose.
It should stop pretending to remain human-readable, go binary, and become a variant of protocol buffers, thrift, etc.
BTW being schemaless is a boon during initial hacking things together, and an impediment when operating and developing the system further down the line. It's the data equivalent of dynamic typing vs static typing, interpreted vs compiled in code.
There are lots of ways standard ways to do serialization to binary formats. https://en.wikipedia.org/wiki/Comparison_of_data-serializati...
> BTW being schemaless is a boon during initial hacking things together, and an impediment when operating and developing the system further down the line. It's the data equivalent of dynamic typing vs static typing, interpreted vs compiled in code.
Sounds like XML, which is a whole other discussion :-)
* JCOF allows for streaming serialization if you want, you can just use the key-value object syntax and use inline strings everywhere. You won't get much size improvement over JSON when using a streaming approach, but you can even do a half-and-half approach; maybe you have a set of strings which you know you're using a lot, so you put those strings in the string table, then you do a streaming encode of the rest of your data, with string references to the string table. You could even do the same with object shapes, where you hard-code a couple of object shapes you know you use a lot, and use the JSON-style key-value syntax for other kinds of objects. Basically, you can choose where you want to be on the efficiency vs streaming curve.
* Nested objects aren't difficult; an object can go in any place where a value can go, just like in JSON. This object can be either a "shaped" object where the keys are listed in the object shape table, or it can be a "keyed" object where you write out the keys inline (as either literal strings or references to the string table).
* JCOF is human-writable; you can write a document which looks a whole lot like JSON. This is a valid document: `;;{"people":[{"name":"Bob",Aage":36},{"name":"Alice","age":40}]}`. But optimal JCOF is mostly unreadable. So it's less human friendly than JSON, more human friendly than binary formats; it's a trade-off.
I think JCOF hits enough interesting, unexplored points on enough relevant curves that it has some value. It also ends up being smaller than any of the binary formats I've tried like MessagePack and CBOR (even with the string references extension), while being text-based. It's obviously not going to replace JSON, but it's useful in some cases.
In short: nice work. Might be fun to roll an implementation or two :-)
But because the Minecraft JSON is highly redundant, I also tried it on 84M of fake social graph data. `jcof` gave 44M, `gzip -9` gave 19M and `zstd` gave 18M.
In summary, neat idea, but you're maybe better off with `gzip` or `zstd` in the real world.
And now has updated data showing that sorting your keys lets `lzma` beat `jcof|lzma` at every level. The egress bill has been reduced! All rejoice!
Provided that you don't run into the issue of one end only talking in language X, which does not have JCOF support due to there not being any libraries for it yet. In the GitHub repo, it seems like there is only the JavaScript reference implementation for now.
JSON: 51949 bytes
JCOF: 37480 bytes (0.721x)
JSON ZIPED: 15178 bytes
Zipped json wins again.If it could somehow produce a significant reduction with fewer CPU cycles, there may be a really niche use case, but I don't see those often.
I had devs try to "optimise" the JSON our services were returning by doing things like abbreviating field names that were repeated very often or replacing long string constants with short integers or removing formatting, and so on. When I asked them to do actual test when the service returns gzipped output and they found out that even reducing the JSON size this way by half does not make any measurable difference on the compressed stream in most cases.
In general text formats are a compromise. They are bulky and and inefficient to parse. What you get in exchange is ease of development.
If you are willing to complicate the format just to improve performance, at some point you no longer get ease of development and you should just switch to binary altogether.
But "easy" wins so we keep returning to CSV and JSON. ^_^
OTOH zstd [1] should be significantly more efficient, while being as fast or faster.
This format is also not interoperable without a decoding library, which kind of invalidates the comparison to JSON. If you're going to do this, why not just go full binary and save even more memory?
For those unfamiliar, here is the interesting bit:
> rkyv implements total zero-copy deserialization, which guarantees that no data is copied during deserialization and no work is done to deserialize data. It achieves this by structuring its encoded representation so that it is the same as the in-memory representation of the source type.
It is super cool that it serializes hash tables and b-trees because that is where serde’s zero copy parsing ends
It seems perfect for my use case, which is putting stuff in a chunk of memory shared between processes
There are a pure-JS DEFLATE implementations with reasonable performance-- I recommend fflate.
For those who disagree, try finding a json library that can reliably round-trip the data in a json file (i.e. ignoring spacing, tabs, and extra line breaks)
Common breaking points are discrimination between integers and doubles (e.g read 0.0 and write 0) and limits to the number of digits in a number.
It’s futile to try and change it now, but I think json would have been better if it had disjunct sets of integers and floats, did 64-bit ints (1) and IEEE doubles (1) as opposed to arbitrary length integers and floats, standardized serialization of time stamps and durations.
Requiring parsers to keep keys in file order also might have been a good thing.
(1) I know that opens a rat’s nest of users wanting unsigned 64-bit integers, signed 16-bit, 128-bit quad floats, etc, but I don’t see how not addressing that problem at all is better than picking reasonable values, event though that gives up on supporting those alternatives. Now, we have libraries making choices there, hopefullly in compatible ways.
We used this to keep our memory footprint low; instead of loading BSON files we would memmap them (getting a pointer that can be used to read the data right off the disk) and passed that to the BSON library. Bang, we could map a 10MB file and it wouldn't use 10MB of system memory; it would just swap portions in and out as they were used.
This has performance tradeoffs but for us (on mobile, where there is no swap space) the memory savings was worth it.
If I need a more compact no longer at all human readable format I can use a compact binary format. If not I can use JSON which has wide spread support.
or just use a compression (after which the difference between JSON and JCOF should be negligible)
Do you mean you'd use a format with a schema?
I think this is a great idea. JSON is incredibly wasteful and any improvement is welcome.
And yes, you can gzip JSON. still wastes a lot of space when you actually need it in a format where you might want to read individual fields.
I love me some jq, not moving much data as JSON, and so this amounts to an interesting exercise.
curl https://raw.githubusercontent.com/mortie/jcof/main/tests/corpus/meteorites.json | gzip -9 | wc
Gives me 34569
So the comparison is: JSON: 244920 bytes
JCOF: 87028 bytes
GZIP: 34569 bytesAuthor mentions gzip doesn’t work for some use cases. For use case mentioned I’d expect sqlite to be similar, at least that is the default thing I’d reach for.
If for some reason sqlite wasn’t sufficient probably a custom binary encoding controlled and updated via code instead of config would be next.
json jcof jcof/json
plain 244975 87083 0.355
gzip -6 35829 33152 0.925
gzip -9 34384 32875 0.956
xz -9 27864 28696 1.030
And I imagine this is close to ideal for jcof. So unless that last few % really matters, gzipped JSON is probably much better in the general case.Just tested my 84M fake social media file - `jcof` gives 44M, `gzip` gives 19M, `jcof+gzip` gives 17M. In essence, you've gained 2M for two CPU intensive procedures instead of one. Doesn't seem all that worth it?
Prompted me to check if the higher zstd levels worked any better on my 84MB fake social graph - nope - and then if LZMA was any good - yes, `lzma` at 5 or higher on the raw JSON beats `jcof | lzma` by ~2M every time. `lzma -4` beats it by ~400k.
If I sort my object keys (a la `jq -S`), `lzma` beats `jcof|lzma` at every level (`gzip` never gets close, `zstd` gets closer.)
They're improvements, they're experiments, they're an attempt to move forward. Somebody cares!
but Randall Munroe said it pretty well: https://xkcd.com/927/