Bebop: An Efficient, Schema-Based Binary Serialization Format
rainway.com
rainway.com
A link to the benchmark code and description of the data would be nice.
It also doesn't show data for FlatBuffer, which is often a lot faster and leaner than ProtoBuf, or for Capt'n Proto.
ProtoBuf is not exactly known for amazing performance or very optimal client implementations across the various languages.
By all means, create a new serialization format, why not. But with so many options to chose from, I would require really strong justification internally.
But yes, it's a lot of work and they probably should document what the trade-offs are that they have made.
I would guess the high speed is the triad of
length encoding with a header vs searching for delimiters
using what they call structs for benchmarks (no repeatedly sending the field name)
how much you trade off safety/sanity checks for performance
Oh, and keeping ints little endian.- The benchmark code is present in the laboratory directory of the repository.
- We don’t compare to Capt’n Proto because it does not have a stable web-based implementation, at least not one that has the features that make it so fast natively, so there is nothing to compare.
- Flatbuffers are fast but have a notoriously awful API to work with while also creating their own non-standard data structures in languages like C++. Bebop generates standard type-safe code.
- Bebop doesn’t try to compress data other than strings. This is because we don’t want to be responsible for compressing trailing zeroes when faster compression algorithms exist that can be down after encoding. Also most data is tiny.
- Bebop supports discriminated unions and has a much more robust type system than Flatbuffers.
- We’re not convincing anyone to use our stuff. It was made for us and open sourced because it was useful; we don’t need people ripping out their current serializers if there’s no pressure to do so.
What would be helpful is a concrete example showing what was tried with an existing approach that fell short. I mean code, benchmarks, theory.
Soo... cap'n proto?
1: https://cs.stackexchange.com/questions/129904/does-there-exi...
I miss "tagged unions" or enums with values a.k.a. sumtypes.
I wish there was a half way house between something like this and JSON as I really find is useful to be able to debug over the wire with Postman or Charles for example.
All other serialization formats using an IDL seem strictly better.
A weird criticism considering the encoding rules usually used for ASN.1 are all binary and some of them are bit-packed (like PER), which is very uncommon in newer protocols (for good reason).
Oh and there is OER now, which is actually a very reasonable binary encoding.
To be fair though, I am open to the idea of having a separate schema definition language, at least. (And please don't say DDL, it doesn't even come close.)
This thing it looks like uses normal C++ structures under the hood, and if so that's a huge plus.
So okay, FlatBuffers doesn't map its zero-copy philosophy perfectly everywhere -- fact of life. What would you offer then? Which other format and/or library?
* Arrow is columnar, batch-oriented, geared toward high throughput.
* Bebob is record-oriented, similar to Avro, Protobuf or JSON, geared toward low latency.