Q^mSat,3^b:d+s+E,4Fri,3^u:h+k+u,6Thu,3^P:j+
If you are effectively going binary, do it. CBOR or Protobuf or any dozen other binary serializations that would be far more efficient.The author claims this is because of copy and pasting… cool, remind me what BASE64 is again?
Sure, you can encode as JSON, then compress with gzip and then base64 encode. You'll probably end up with something smaller than rx and be extremely safe to copy-paste. But your consumers are going to consume orders of magnitude more CPU reading data from this document.
RX is usable as-is, is compressed, and is copy-pasteable. It's the unique combination of properties that makes it interesting.
>Q^mSat,3^b:d+s+E,4Fri,3^u:h+k+u,6Thu,3^P:j+
My man… no. I have no doubt you could kind of figure out what that sample is hot off the heels of writing this, and likely not in six months. And to consider that anyone else would fill their brain with the rules to decipher that, Nah 2.0.
Even a trivial doc like this is challenging for me to read as a human.
- clipboards
- logs
- terminal output
- alerts
- yaml configs
- JSON configs
- hacker news comments
- markdown documentation
- etc...
I assure you, this is not a solution looking for a problem. I started out with binary encodings first, but then realized it limits so many workflows.
> Nah.com, fam.
Being able to copy/paste a serialization format is not really a feature i think i would care about.
> None of the space savings/efficiency of binary
For string heavy datasets, it's nearly the same encoding size as binary. I get 18x smaller sizes compared to JSON for my production datasets. This was originally designed as a binary format years ago (https://github.com/creationix/nibs) and then later after several iterations, converted to text.
> Being able to copy/paste a serialization format is not really a feature i think i would care about
Imagine being paged at 3am because some cache in some remote server got poisoned with a bad value (unrelated to the format itself). You load the value in dashboard, but it's encoded as CBOR or some binary format and so you have to download it in a binary safe way, upload that binary file to some tooling or install a cbor reader to your CLI. But then you realize that you don't have exec access to the k8s pods for security reasons, but do have access to a web-based terminal. Again, to extract a binary value you would need to create a shell, hexdump the file and somehow copy-paste that huge hexdump from the web-based terminal to your local machine, un-hex dump it, and finally load it into some CBOR reader.
A text format, however is as simple as copy-paste the value from the dashboard and paste into some online tool like https://rx.run/ to view the contents.
Unless, to read that correctly, it only has a text encoding as long as you can guarantee you don't have any unicode?
I did add some small examples to the repo.
https://github.com/creationix/rx/blob/main/samples/quest-log...
The older, slightly outdated, design spec is in the older rex repo (this format was spun out of the rex project when I realized it's actually a good standalone format)
https://github.com/creationix/rex/blob/main/rexc-bytecode.md
Oof.
The format is technically a binary format in that length prefixes are counts of bytes. But in practice it is a textual format since you can almost always copy-paste RX values from logs to chat messages to web forms without breaking it.
unciode doesn't break anything since strings are encoded as raw unicode with utf-8 byte length prefixes. It supports unicode perfectly.
If your data only contains 7-bit ASCII strings, the entire encoding is ASCII. If your data contains unicode, RX won't escape it, so the final encoding will contain unicode as UTF-8.
No human is reading much data regardless of the format.
What is the benefit over using for example BSON?
If all your workflows allow copying as binary files, more power to you! But there are a lot of workflows where that is not possible. This was inspired by years of hands-on operational incident handling in production systems. Every time we use a binary format, it's extra painful.
This particular format would be slightly more compact as binary, but not enough to justify closing the door on all the use cases that would preclude.
I'll probably add a binary variant for people who prefer that (or for people who want to be able to embed binary values in the data without base64 encoding it)
another thing, I put in a 400KB json and the REXC is 250KB, cool, but ideally the viewer should also tell me the compressed sizes, because that same json is 65kb after zstd, no idea how well your REXC will compress
edit: I think I figured out you can right click "copy as REXC" on the top object in the viewer to get an output, and compressed it, same document as my json compressed to 110kb, so this is not great... 2x the size of json after compression.
The primary use case is not compression, it's just a nice side effect of the deduplication. This will never beat something like zstd, brotli, or even gzip.
My production use cases are unique in that I can't afford the CPU to decompress to JSON and then parse to native objects. But with this format, I can use the text as-is with zero preprocessing and as a bonus my datasets are 18x smaller.
Right and that makes sense. There is more information in here. The entire thing is length prefixed and even indexed for O(1) array lookups and O(log2 N) object lookups.
If you don't care about random access and you don't mind the overhead of decompression, don't use RX.
Let me know what you think
https://github.com/creationix/rx/blob/main/README.md#when-to...
Yes, serially. Which means no random-access across the transfer channel.
Yes, many formats are read start-to-end, but I don't think that's a requirement. The important thing is it can be stored and transmitted as a stream of bytes. The word describes how it is transported and stored, not how it is read.
Yes. That's precisely what "random access" means.
> The important thing is it can be stored and transmitted as a stream of bytes.
What isn't stored and transmitted as a stream of bytes? Memory itself is a sequential array of bytes. The criteria you're using here seem to be all-encompassing.
When we're talking about serial access to data, it means that we're sequentially parsing the stream of bytes from its starting point, rather than reading arbitrarily from any point in the stream that we desire.
Yes, by your definition, this is random and not serially read.
(Or to avoid using cat to read, whatever2json file.whatever | jq)
What might be interesting is to have a tool that processes full JSON data and creates a b-tree index on specified keys. Then you could run searches against the index that return byte offsets you can use for actual random access on the original JSON.
OTOH, this is basically just recreating a database, just using raw JSON as its storage format.
Maybe a query over the random-access file then converted into JSON would work?
I did build that once. But keeping track of the index is a pain. Sometimes I was able to generate the index on-demand and cache it in some ephemeral storage, but overall it didn't work out so well.
This system with RX will work better because I get the indexes built-in to the data file and can always convert it back to JSON if needed.