I am not 100% sure why I would choose to use this - maybe for super tiny documents?
I am not 100% sure why I would choose to use this - maybe for super tiny documents?
{"red": 127,"green": 28,"blue": 27,"x": -18,"y": 27,"z": 2008,"time": 1583862185,"dollars": 1298,"cents":12,"name":"tom",timeseries":[1,2,3,4,55,66,77,218,239,340,9009,80008,400004]}
gives 106 bytes with msgpack and 157 with gzip. Also, if you gzip the msgpack, for your message example, it's 248 bytes.
- https://docs.microsoft.com/en-us/security-updates/securitybu... - https://en.wikipedia.org/wiki/Billion_laughs_attack - https://en.wikipedia.org/wiki/Zip_bomb - ...
current_addr + message_len - start_addr < buffer_len
Or am I missing something?Similar to HeartBleed, where there wasn't validation on the heartbeat message, and the server would echo back buffer_len instead of just what was sent.
I can think of a very contrived situation where this can be a problem, but in most cases this will be perfectly safe.
ie just because MessagePack is a binary format it doesn't mean you can skip the same string checks that JSON requires; which means parsing MessagePack strings is unlikely to be any faster than JSON strings (contrary to the suggestions others have implied with the "just memcpy" comments). It's just with JSON that validation is done as part of the parser (remember JSON only technically supports a subset of ASCII and any extended characters or unicode is encoded via escape codes) where as with MessagePack you'd need to do that validation as an additional step.
Integers, on the hand, might differ since JSON would need additional validation (again, backed into the parser) which MessagePack would not because MessagePack encodes the integers as binary integers where as JSON encodes them as ASCII values that would need converting back to binary integers.
(hint: read the message I'm replying to).
However if you’re accepting MessagePack encoded data from insecure systems (such as end users) then you absolutely should be validating your input somewhere along the pipeline and it’s usually better to do that early on.
Also it’s not generally the distributed systems you worry about when it comes to this specific degree of micro-optimisation (which is basically what this is). It’s the monolithic ones. Distributed architecture is meant to solve various problems (for example but not limited to, high availability, reduced geographical latency, single site but running on cheaper commodity hardware, etc) but often at the cost of CPU cycles. Whereas your monolithic infrastructures where you have fewer servers (such as Stack Overflows set up) would be greatly more dependant on reducing computational overhead where corners could be cut. However they’d also be significantly less likely to need networked RPCs via MessagePack anyway (simply due to the monolithic design of their architecture).
Yes, so what's the problem?
Meanwhile keep in mind that the JSON document was already passed in a HTTP request body, or that its trivial to put string length checks on a parser, or even nesting limit checks.
That said, not something I want to prematurely optimize for.
I don't see your point at all. The state machine required to parse a JSON string essentially has only 2 states: current token is either a string character or a string delimiter. That's it. Adding a record size increases the number of states because now you not only need to track if a token is a string character but also if the character makes sense to be there within that state.
Moreover, when you parse a JSON string, once you hit the string's end delimiter you already are able to calculate the string size. Thus, if the string is short enough to fit the buffer then not only is the lexer simpler but it also requires the same memory allocations than strig formats with record sizes. If however a string doesn't fit the buffer then we are already in the territory of a ropes data structure in both cases, thus the number of memory allocations tend to be equivalent.
Without GZIP:
JSON 583 bytes
MessagePack 304 bytes
52 %
With GZIP:
JSON (GZIP) 289 bytes
MessagePack (GZIP) 248 bytes
85%
I just took a quick look and they do compress their queries, using Brotli, which is even more efficient than GZIP. They are behind Cloudflare, so it's probably from them.
I just tried with https://jsonplaceholder.typicode.com/todos and I get this:
Raw:
JSON: 24,311
MessagePack: 14,704
60 %
With GZIP:
JSON: 3,965
MessagePack: 4,063
102%
With Brotli:
JSON: 3,495
MessagePack: 3,704
106%
So essentially, MessagePack compressed is WORSE than JSON, and not compressed, is at least 3 times worse than compressed.
In the case where your data can't be cached but CloudFlare proxies and compresses for you, MessagePack wins because your uncompressed connection to CF is 40% more bandwidth efficient so egress out of your app server for any traffic is reduced. In the case where CF can cache responses you get the same back end bandwidth win for a negligible (if any actual) amount of extra egress out of CF.
Use what you want but if you need a schemaless serialization format, MessagePack works well and is compact for "free".
I guess the ability to have a schemalass and no type checking format (just like JSON), while enjoying some performance benefits? It just feels like a weird niche to me.
Why would it be your default?
https://github.com/thekvs/cpp-serializers#results
That said, currently I don't work with it in any projects. Apparently it hasn't really caught on.
At the end of the day, though, I'd seriously much rather just have a halfway well commented *.proto file. Work smarter, not harder. It's just that I won't recommend that for public-facing APIs because GRPC isn't universally well supported, while JSON over HTTP is.
Avro is pretty similar to MsgPack in concept. It seems less popular (at least in my circles), and has fewer implementations, which may or may not matter. As for performance and efficiency, the first benchmark I found [1] shows MsgPack is more space efficient, and faster to serialize, while Avro is faster to de-serialize. The second benchmark I found [2] found exactly the opposite. So it's not clear to me that either is better, and I'm sure it depends on your data, which library you're using, and how you're using it.
Being nearly as performant as the predefined-schema serializers, while being nearly a drop-in replacement for JSON, seems like a major and valuable use case to me.
[1]: https://medium.com/@nitinpaliwal87/compression-and-serializa... [2]: https://github.com/saint1991/serialization-benchmark
While I see the same numbers as you do, there's no way a format that includes plaintext keys alongside the values is going to be smaller than a format that only includes a packed field number+type and the corresponding values. There's something fishy going on here, and I can't seem to find a link to the source behind his benchmarks?
In force documentation:
https://www.itu.int/rec/dologin_pub.asp?lang=e&id=T-REC-X.68...
Blog on ASN.1 schema for JSON:
The big difference is that messagepack is schemaless.
>ASN.1 is a mature standard. As I already mentioned, it has been around since the 1980’s. Though stable, it is not stagnant; the most recent revision occurred in 2008.
It might be old but it is not widely used. To me it doesn't matter if something was created yesterday or 20 years ago as long as I can use it and using ASN.1 is significantly harder than necessary.
>ASN.1 has tool support. There are both commercial and open source tools that will generate code from your ASN.1 specification.
Ok, so where is it and why is it so difficult to find?
Even if we assume that we should just build it today then all the above claims suddenly become worthless. If someone builds ASN.1 tooling in 2020 then it is just as immature as e.g. GraphQL tooling that was built 2020. If there is renewed interest in ASN.1 then it might gain new features that will cause it to become less mature/stable again.
Having "mature" software is useless if it doesn't meet user demands.
I've spent a lot of hours poring over hex dumps of BER messages and CBOR messages (which are basically MsgPack) and I vastly prefer CBOR. But I prefer JSON way more.
How much time would it take you to either
a) turn on gzip support on your HTTP server and keep using JSON
b) rip out your JSON stuff from controllers and whatnot and replace it with a custom serializer/deserializer that's far from an industry standard?
Edit: found my post https://stackoverflow.com/questions/9884080/fastest-packing-...
I no longer have any connection to this repo, but the benchmarks (and a schema validator gem I wrote) are here: https://github.com/deseretbook/classy_hash/blob/master/READM...
With MessagePack the schema is embedded alongside the data, exactly like JSON. So it's just binary instead of text encoding, which saves some space, but as you pointed out, standard text compression algorithms are going to likely perform similar or superior.
B64 as originally used (with an extra >2.6% for line splitting) is handily over 35%. As we use it now it’s still a bit over 33%, because of the == endings to detect truncation... 66% of the time.
1. msgpack + gz would be a more fair comparison to json + gz when comparing file size
2. have you run any time comparison? It would seem to me that msgpack gets pretty close to json + gz with far fewer resources, time
I havent had need to deal with such constraints personally, but maybe its something I should consider more. It doesnt need to be a 10x win to still be useful, and if its a simple import to use then i would likely be crazy not to consider it.
Typical gzip decompression speed is somewhere in the 200-250 MB/s region, compression is much slower. LZ4 for example tends to compress at ~600-700 MB/s, and decompress at several GB/s. zstd is tweakable over a very wide range of ratio-speed trade-offs.
LZMA(2) (xz) is a rather troubled format and should not be used any more. bzip2 has always been slower than gzip with usually marginally better compression. It has been irrelevant for a long time.
You mean xz. Not sure where you got the idea that it’s a troubled format, but if you’re talking about the infamous “Xz format inadequate for long-term archiving”, IMO that’s just bzip2 authors taking a dump on xz for no good reason, and fortunately for us it’s bzip2 that’s basically irrelevant today, not xz.
Funny you say that, my company uses bz2 for compressing pretty much everything.
Brotli is better, zstd is faster, lz4 is often good enough and a lot faster.
gzip is fast relative to other compression algorithms and relative to internet speeds.
gzip is slower than an ideal gigabit network, but faster than 100mb.
https://www.rootusers.com/gzip-vs-bzip2-vs-xz-performance-co...
> gzip is fast relative to other compression algorithms
gzip looks fast perhaps compared to "xz -9", but not to anything modern.
Gzip is quite close to the Pareto frontier, meaning it is a good trade off of time and space. See the charts at http://mattmahoney.net/dc/text.html (And read up on the Hutter Prize!)