All in all, it's something to study and learn from, but I strongly recommend not becoming involved unless you are 100% happy with being tied into node.js.
All in all, it's something to study and learn from, but I strongly recommend not becoming involved unless you are 100% happy with being tied into node.js.
To me it's worse that key order must be preserved, which this standard does not specify the way I understand it.
Source: https://ssbc.github.io/scuttlebutt-protocol-guide/#message-f...
- Two spaces for indentation.
- Dictionary entries and list elements each on their own line.
- etc...
This is so weird to me. Why a protocol needs a strict, opinionated format of JSON? If they really need a very specific format of JSON, why they even bother JSON? There are better options like protobuf.
This seems the worst example of "Use JSON for everything".
If they wanted both the descriptions to be visible and the order to be preserved, they could use:
[ ["previous", "..."], ["author", "..."], ...Now, the feed format is indeed a practical problem if you try to make a client from scratch and it is annoying. Luckily, using one of the already existing libraries will handle this for you.
Changing that feed format for everyone is not possible, simply because there's already an existing social network built on the old ones that we very much want to preserve since we actually... well... hang out there. Changing feed format thus involves adding a new feed format and making sure other clients can handle and link together both. At the benefit of abstracting away the feed format, and being able to iterate on them.
This is the exact reason bencode (https://en.wikipedia.org/wiki/Bencode) was invented, and I still believe we can replace all uses of json by bencode and be better off it, because it solves all too common issues:
- bencoding maps are in lexicographical order of the keys, so no confusion possible for hashing/signing (a torrent id is the hash of a bencoding map)
- bencoding is binary friendly, in fact it must be because it stores the pieces hashes of the torrent
Why don't we use bencoding everywhere ?
* It is very easy to parse/produce Bencode
* It's probably fast
* The specification is really bad
* No float type
* No string type, just byte sequences. This is especially bad because most/all dictionary keys will be utf-8 strings in practice but you can't rely on it
* Integers are arbitrary length, most implementations just ignore this
I think Bencode is an ok format for its use case, but I don't think it should be used instead of json.
I thought about sorting keys and other things like that, and the dozen edge cases and potential malleability issues dissuaded me for the compatibility issues mentioned above.
How have others solved it?
As I put in another comment, a torrent id is a hash of a map, where one of the keys contains binary data. bencoding solved that decades ago already.
BON is compatible with json+ and erlang data type, in specific, it allows any data type for the map key. Json only allows string as map key.
Which protocols? Can you point to examples?
Because I was about to design a signing Json solution but based on the comments here it is a bad idea.
1. Canonicalize and sign: the format has a defined canonical form. Convert to that before doing cryptographic operations. If the format is well designed around it, this is doable, whereas JSON doesn't really have this and with many libraries it's hard to control the output to the degree that you'd need. 2. Serialize and sign: serialize the data as flat bytes, and then your signed message just has a "bytes" field that is the signed object. This is conceptually not far off from the base64 solution above, except that there's not extra overhead, since with a binary protocol you'll have a length prefix instead of having to escape stuff.
Personally I'll just concatenate the values in a defined order and sign/hash that.
A simple proposal for actual security guys to rip to shreds:
Strings are represented as utf-8 blobs, and hashed.
Numbers should probably be represented using a proper decimal format and then hashed. If you're reading the JSON in and converting it to floats, you could get slight disagreement in some cases.
Arrays are a list of hashes, which itself is hashed.
For objects, convert the keys to utf-8 and append the value (which will always be the hashed representation) and then sort these entries bitwise. And then hash all that.
Or, better, it'd be great to have an order-independent hash function that's also secure. I doubt xoring all the pairs would be good enough. Update: a possible technique[1]
[1]: https://crypto.stackexchange.com/questions/51258/is-there-su...
The protocol helper functions would be compiled to a webassembly library, and you would reuse them in Go, Python, the browser, etc
Of course, it's not justification for using their protocol (rewriting a protocol in another language is a good test for the protocol specification), but that would be a usecase for webassembly.
Just giving people the heads-up that JS seemingly is the only blessed language for scuttlebutt
Rust:
https://github.com/sunrise-choir/ssb-legacy-msg/blob/master/...
https://github.com/sunrise-choir/ssb-publish/blob/master/src...
Go:
https://github.com/cryptoscope/ssb/blob/master/message/legac...
https://github.com/cryptoscope/ssb/blob/master/message/legac...
Spec:
https://spec.scuttlebutt.nz/feed/messages.html#json-encoding
https://spec.scuttlebutt.nz/feed/datamodel.html#signing-enco...
But that isn’t a reason why scuttlebutt isn’t more popular. It only takes one Go implementation and the problem is solved permanently.