CBOR – Concise Binary Object Representation
cbor.io
cbor.io
How does the density compare with JSON run through GZIP compression?
I'm not sure what 'ambiguous decoding' you're referring to. I've seen some complains that an int might decode to an int32 or int64 depending on the value but be different than the original storage size. e.g. '42' stored from an int64 might unpack to an int8 depending on the language and implementation of the decoder.
EDIT: Found bigints. Still weird having a signed 65-bit integer.
ioquatix: You might want to edit your post to indicate exponentiation. It took me quite a moment to realize you meant that a signed 64-bit int represented in 2s complement has two-to-the-sixty-fourth values, not two hundred and sixty four. I'm guessing you used two asterisks which HN mistook as a zero-length italics?
If you happen to be on OS X, you can type Control+Command+Space in most contexts, and get a popup that will let you select special characters, like the above, emoji, etc. It even has a decent search box.
If you're on Linux, you can set a key as a Compose key (I use my right alt key); the default set of compose sequences includes superscript six as Compose, ^, 6. You can also input arbitrary unicode into most contexts if the app is using GTK with Ctrl+Shift+u, followed by the hexadecimal code point, followed by space.
On Windows, I have no idea.
[1]: http://www.fileformat.info/info/unicode/char/2076/index.htm
Having written an implementation (the Lua one) I quite like this format, it's flexible, not too complicated and still quite compact.
It also seems to have a better implementation of streaming data than msgpack. In msgpack, streaming is something you have to implement outside of the format itself, possibly by concatenating together many msgpack representations. CBOR has a way to say "here comes a streaming list, I'll tell you when it's done".
BSON is a representation of MongoDB's data model and doesn't make that much sense to use without MongoDB.
I currently use msgpack for a lot of things, but if CBOR's Python library is good enough, I might switch.
I read somewhere that CBOR was better designed for extensibility, but don't know anything further about it.
One difference (on the non-technical side) is that CBOR is standardized through IETF.
The problem is that the 'str' type contains arbitrary binary data in an unspecified encoding, and always will, because of backward compatibility. This isn't changed by adding a 'bin' type.
Msgpack decoders in Python, for example, have to give you bytestrings unless you pass an option that promises that 'str's are all encoded in UTF-8.
Raw
String extending Raw type represents a UTF-8 string
Binary extending Raw type represents a byte arrayBut the link you posted referred to three types, where "str" and "bin" subclass "raw", which sounded like it provided a non-backward-compatible "str" that's guaranteed to be text.
Tags are used to apply meta information to a piece of datum. For example, you can tag a UTF-8 string as a URL or tag an array with a reference so it can be referred to elsewhere in the CBOR encoded data (an extension defined by the IETF but outside RFC-7049).
An interesting feature is support for strongly typed arrays and objects. This lets you embed binary data as-is by specifying it as an array of uint8.
when are we going to learn from our past mistakes?
Contrast to Rust's Option<Foo> type; you know when null (None in Rust) is a possibility, because it's in the type. (And in Rust, you're forced to deal with it, too.)
But they cause problems even without static types.
https://www.lucidchart.com/techblog/2015/08/31/the-worst-mis...
For example, Rust can have Option<Option<T>>, which can be None, Some(None), or Some(Some(...)). But you can represent that with null, because nulls don't "stack".
ASN.1 would perhaps be the main motivation here. Despite being an abstract syntax, signatures across BER/DER/PER were all distinct. With a little bit of work, signature algorithms as expressed in standards like CMS could be abstract across the encoded representations.
But They Didn't Do That.
Flash forward almost 30 years into the future, and we're literally dealing with the same problems:
https://www.ietf.org/proceedings/94/cose.html
"The resulting formats will not be cryptographically convertible from or to JOSE formats."
WHAT REALLY? A longstanding problem with these formats for 30 years, and they literally did the exact same thing? Yes, yes they did.
Hey, know what format does this (at least for bearer credentials if you think JWT/CWT are cool. It's hard to argue with one specific JOSE/COSE thing since this cancer literally has their fingers in every single honeypot they can get their tentacles into)?
Macaroons:
In addition to not just being JSON-for-CMS/SAML, Macaroons are actually designed to be bearer credentials rather than slapping a bunch of "We took an old idea and added JSON" and slapping it on old concepts, but...
Macaroons are provably secure in their own dialect of Abadi's authorization logic. Check the last page of the paper:
http://research.google.com/pubs/pub41892.html
CWT (and vicariously JWT) try to slap JSON/JOSE syntax/standards onto a bunch of existing concepts, but fail to actually fix the fundamental problems, like provable security. Yes, provable security: Macaroons are predicated on proofs. JWTs are predicated on broken promises and ad hoc design.
All that said, the authors of the JOSE/COSE standards couldn't have tried harder to repeat every single one of the mistakes of the past. These standards are nonsense. Unless you're switching from CMS they offer no practical benefits, and can potentially introduce new vulnerabilities due to their ludicrous complexity.
Unless you can name the exact CBOR-encoded standard you want to use and why you should use it, and the alternative you're considering is practically anything but CMS, AVOID AVOID. Stick with anything that's more standard. Slapping JSON on things doesn't help, but just introduces new problems in a space where older standards are at least better understood.
There are problems in this space: ASN.1 is old, overcomplicated, hard to use, and the source of many vulnerabilities in the way it's described. So we have a serialization format with bad semantics used in security critical contexts. Should we replace it? Yes? Is JOSE/COSE the solution?
JSON has many odd/ambiguous semantics, a shitty type system, and there is no direct mapping between ASN.1's type system and JSON/JOSE's, because it is intentionally restricted by design.
JOSE/COSE's solution is to improve a security critical format by getting rid of types.
Wait what? Improving security by getting rid of types? Yes that's exactly what JOSE/COSE are doing, and it's the wrong direction IMO.
I would much prefer some sort of modern typed serialization format which has packed and unpacked representations. Protobufs or capnproto come to mind.