Saltpack – A modern crypto messaging format
saltpack.org
saltpack.org
Messagepack is schemaless and noncanonical. What that means is that a lot of the bounds/field checking is pushed up to the application layer. I wouldn't encode crypto with that (and I love Messagepack).
All the hate for ASN.1, yet it is among the most battle-tested specifications out there. Blaming ASN.1 for the shitty ASN.1 parsers written in the 80's and 90's is like blaming libsocket for all the network attacks.
"CBOR is better than MessagePack because it's better designed" is a good argument, I think --- but: you'd want it to be so much better that the difference is material.
"CBOR is better than MessagePack because it's standardized" is not, I think, a persuasive argument.
I agree that IETF standardization isn't a good argument (well, for some it is), that's why I replied with a better argument :) But, seriously, I won't say that everyone should replace MessagePack everywhere with CBOR, both work fine (as long as you use the latest version of MessagePack, with binary/string distinction).
The tags are the worst, actually. Sure, spec says "decoders do not need to understand tags", but this is not really the case. For example, if someone has floating point numbers and worries about precision loss, they can store the value as _decimal fraction_ (per section 2.4.3). This means that your decoders have to support both tagging and your favorite bignum library just to make sense of the data. In comparison, in msgpack (or json or xml or anything else) you would just have to store a string representation -- trivial to convert to either regular floats if you do not care, or to pass to your favorite bignum library (and this will be simple, as all of them support constructor based on ascii strings).
In general, I think optional tags in data-interchange protocols are a very bad idea. For example, there is a tag for "Standard date/time string" and for "Epoch-based date/time". Which means that either:
- You schema says "date/time", and your decoder now must support both of them (and probably untagged strings, and integers too). So this is an extra complexity in your decoder.
- Your schema says "date/time in 'Standard date/time string' format", and now every encoder user must make sure the emitted value is encoded and tagged appropriately. This means you cannot do `x = cbor_encode({"now": date})`, you have to read your encoder documentation to make sure the CBOR encoder you are using will generate the required encoding.
So extra complexity in either case, and no real benefits. Better stick to msgpack, at least it has no extensions defined currently.
Yes, you have to read the documentation of your encoder/decoder to understand what tag values it maps to your programming language's objects, but if you need to encode or decode those same values with MessagePack you'll have to define your own format for them and document it. You just moved this problem up the stack, but with an ad-hoc format.
Separation of "null" and "undefined" is for full JavaScript support. Before C99, C didn't have boolean type, but you wouldn't complain if serialization formats had them, would you? Same thing "undefined": while it's useless in most other languages, it's useful to have it for JavaScript.
I don't like float16 too - they took 19 lines of decoder code to support (no need to support it in encoder) — but it's the same situation as with "undefined" — some people need it.
https://github.com/msgpack/msgpack/blob/master/spec.md#forma...
I experimented using them to implement a limited atom/token spec for use on Arduino.
Disclaimer: I wrote and maintain a MessagePack library.
Tagged values are useful! Because of the simplicity of CBOR (compared to MessagePack), it's about 3 lines of code to implement them in the encoder.
I'm not sure I understand your complaint about canonical mode. First of all, it's optional, and secondly, it's actually good that it's defined in the spec, otherwise you'd have multiple incompatible formats like canonical JSON.
The point about strictness and canonical mode are that they are poorly thought out. For example, sorting based on binary representation requires multi-pass encoding which is slow, complex, error prone, and completely non-intuitive: [1,2,3] comes before 100000 which comes before [1,2,3,4].
This is very very bad.
ASN.1 (which is also used in x509) is a catastrophicly awful format for encoding certificates or other crypto operations, and is the direct cause of a huge amount of SSL/TLS security issues.
ASN.1 is a horror show, and the actual implementations of it that are available are even worse-- I've still yet to find an open source implementation of BER that strictly matches the spec, correctly accepting and interpreting all values that should be accepted and rejecting all values that should not... but I believe that you're joking about "turing complete".
... But ASN.1 is so bad that I'm not completely sure. Can you confirm or cite?
2. CMS requires certificates, which are harder to generate, verify and deal with than simple public keys.
3. CMS allows you to use any combination of ciphers, including weak one. Experience with TLS told us that this flexibility is NOT good feature when it comes to crypto standards.
Besides that, all the reasons the rationale raised against PGP also reamin true for CMS: https://saltpack.org/pgp-message-format-problems
Actually, yeah, that's pretty much what they said:
The changes here are small: we've reduced our characters to base62 plus some period markers, and only at the ends of words. PGP messages often get mangled by different apps, websites, and smart text processors.
You could argue that this doesn't need to be part of a message format itself because it'd be perfectly possible to write an intermediation layer for print that simply translates arbitrary input to an OCR favorable output and back, and maybe that'd make sense anyway if there are other desirable choices for properties like parity unique to print/archival/OCR. Still, to the extent that a format base choice is arbitrary and makes no particular difference to the humans or computers involved since they'll be intermediating through software anyway, better OCR properties doesn't seem like an entirely unreasonable metric to consider as part of the design considerations if there isn't a compelling reason otherwise.
Also QR codes hit a certain practicality limit with size (2953 bytes to stay in spec)
Perhaps a meta-point is that when you are trying to design a general purpose interchange format, there will always be scenarios that you didn't imagine. In this case I have raised OCR and (legitimately) many people's responses have been a rather polite WTF (although I did garner one downvote). Experience teaches us that formats will be used in unexpected ways.
It also occurs to me that OCR implies a computer, but it could just as legitimately mean a human reading text
Round two of skepticism: msgpack is a niche player with no clear big corporate sponsor. Protobufs, flatbufs, and thrift are all actively making faster better quicker implementations, but I can not off the top of my head think of any major msgpack lovers. Avro also seems to just generally have some fast impls already, especially on platforms I care about[1], so credit there too. I ought review, but out of hand I can't think of anything distinguishing about msgpack.
Definitely nice having some alternative to Salmon protocol[2] (as in Buzz, OStatus) on hand. Alas I believe it's again fully encapsulating, versus say http signatures[3], where the signature is decoupled from the payload. It takes both types!! Neither is right.
[1] https://github.com/mtth/avsc/wiki/Benchmarks [2] http://www.salmon-protocol.org/ [3] https://tools.ietf.org/html/draft-cavage-http-signatures-06
> Not totally sure what BaseX is, how it compares versus Base64, especially post HTTP deflate compressions, but I'm not sure I like it.
BaseX is "armor" -- a way to insert binary data into the text-only media such as email messages and forum posts. It seems to be way more convenient, as it uses no punctuation and is whitespace-insensetive. For example, I have once tries to send PGP message via gmail web interface and it took me two or three tries to figure out how to properly paste the text without gmail inserting extra whitespace and making the message undecodeable. BaseX should not have such problems. It is slightly less efficient that base64, including under compression, but I think it is worth it.
> I'm pretty sure this kind of exercise is better left out, and that everyone should just use zstd on whatever encoding so as to decouple problem domains.
Are you talking about zstd as in "Zstandard - Fast real-time compression algorithm"? This has nothing to do with this proposal, there is no compression anywhere in there.
> Round two of skepticism: msgpack is a niche player with no clear big corporate sponsor. Protobufs, flatbufs, and thrift are all actively making faster better quicker implementations, but I can not off the top of my head think of any major msgpack lovers. Avro also seems to just generally have some fast impls already, especially on platforms I care about[1], so credit there too.
The deserialization speed does not really matter. All saltpack messages are encrypted, and your decryption time will dominate your deserialization time.
> I ought review, but out of hand I can't think of anything distinguishing about msgpack.
Well, msgpack is in a completely different group from Protobufs/flatbufs/and thrift. The former is a protocol, which is implemented by a number of libraries, while the latter is a specific library, available from a single vendor only.
As a result, msgpack has fewer features (no schema support at all), but is not bound to a single large corporation. IMHO, a right choice for global communication protocol.
> Definitely nice having some alternative to Salmon protocol[2] (as in Buzz, OStatus) on hand.
Wait, what? Salmon is about blogs on the web, HTTP posts, XML schemas embedded in HTML. The encryption is only a small part of it.
Saltpack is about encryption email/chat/messaging, has no specification of payload format, and designed to work without HTTP using efficient formats. I see very few common things between two protocols.
> Alas I believe it's again fully encapsulating, versus say http signatures[3], where the signature is decoupled from the payload. It takes both types!! Neither is right.
Well, the big difference is that http signatures have no encryption while saltpack has it. So just two very different goals that they want to achieve.
1. You randomly talk about BaseX ad nauseum but brely mention what I was actually critical of- it's advantages vs base64. I still have no clean picture why not base64, like everyone else, which would serve the exact same needs.
2. You just finish talking about compression then slam me for mentioning compression as a relevant factor.
3. You criticize deser speed as not important, say decryption will dominate. But while the message format may be competing with PGP it's inner payload is msgpack and I expect inside the firewall systems to have canonical text as msgpack, and performance is relevant. But I didn't initially grok that msgpack is the inner payload, that the BaseX text is the normal messaging format.
4. You try to pull some technical distinction nonsense about msgpack being just a messaging format by talking about how competitors also have other stuff too, while also being a message format. I find this distinction in bad faith and believe most techs could reasonably see there is way more overlap than differences. Your advantage ends up rather accurately being "msgpcka has fewer features" and chalk up the advantage as some nebulous political one, while ignoring the fact that Thrift is owned by the most reliable open source org on the planet Apache whereas Msgpack is just some rando project.
5. You obviously don't know what Salmon is for. Your decision to focus on some apparently unrelated technical things that it corresponds with ignores that it is a (XML)signing format for arbitrary content. HTTP and HTML have no bearing on what Salmon Digital Signing Protocol is, yet you harp on them to draw a false contrast, and as usual you refuse to acknowledge that there could be some bearing or relationship that I had validly called out.
I'd like to better understand some of your valid points, but some of your arguments seem done in very bad faith and there to argue rather than explain or demonstrate. I have a hard time understanding how I can start to reconcile our two point of views with what you have written.
Unfortunately, I'm not familiarized enough with the subject to comment on how each format stacks up against various alternatives.
I sometimes scroll through an app's licenses page, and I've spotted MessagePack a few times in big-name companies. It's just not a very flashy subject.
Disclaimer: MessagePack author is an ex-coworker.