Open Sourcing the Stupid-Simple Messaging Protocol
aerofs.com
aerofs.com
Netstrings are so brilliantly simple, see the wikipedia page: https://en.wikipedia.org/wiki/Netstring
This is what DJB says about the netstrings [1]:
> The famous Finger security hole may be blamed on Finger's use of the CRLF encoding. In that encoding, each string is simply terminated by CRLF. This encoding has several problems. Most importantly, it does not declare the string size in advance. This means that a correct CRLF parser must be prepared to ask for more and more memory as it is reading the string. In the case of Finger, a lazy implementor found this to be too much trouble; instead he simply declared a fixed-size buffer and used C's gets() function. The rest is history.
> In contrast, as the above sample code shows, it is very easy to handle netstrings without risking buffer overflow. Thus widespread use of netstrings may improve network security.
[1] http://cr.yp.to/proto/netstrings.txt
BTW, see Aaron Swartz's blog post on DJB, http://www.aaronsw.com/weblog/djb
Buffer overflow concerns are not applicable to SSMP however, as the spec explicitly restricts message size to a maximum of 1024 bytes.
Seems reasonable enough?
Please put it in your abnf spec. People use them, you know.
Yes, it's hard to specify. The problem is that you have both an uncapped ID and an uncapped PAYLOAD in the same message. I recommend giving ID a max length of, say, 32, and PAYLOAD then has a max length of 951 if I'm counting right.
Or you could consider that an IPv6 path MTU is at least 1280, and use that (or 1232) as your per message bound instead of 1024. You're sending a packet, might as well get full value.
https://en.wikipedia.org/wiki/String_%28computer_science%29#...
Though I suspect my parsing code still have bugs and exploits.... Looks like I need to put a guard in get_util for the max size.
The CRLF encoding is not so bad in this case: read 8 characters, in parallel detect CR. If there is no CR in your word, just append the data to the buffer. When you do have a CR, it's a big pain: you need to save the last word with byte masks, then shift any remaining for the next input (and the entire next string is shifted by this left-over balance). You could try to make all strings a multiple of 8 in length to avoid this, but this adds overhead to the message so is inefficient- the hardware will just have to do it.
OK, so now in your new format the hardware has to parse a variable length decimal number and convert it to binary (ideally in parallel), very fun! You could make the conversion byte at a time, but it's slow. You need to implement overflow detection.
At the very least use hex instead of decimal. Even in software you may need overflow detection. This is easy in hex, not so much in decimal. Better is to require the number to be a multiple of four or eight digits, even though this is a waste of bandwidth.
In this case, it is a messaging protocol. The incoming message essentially must be copied somewhere, therefore space must be allocated to store it, and therefore the length must be known.
CRLF require either two passes (one to get string length, another to copy data) or continuously expanding storage, both of which are significantly more expensive than just parsing a short number in the beginning of the string.
In hardware there is no malloc- instead there is a pool of pages and the string would be stored as a linked list of such pages. Linux socket buffers do the same thing.
Now, how is that relevant to a chat protocol?
Note: Reminds me when I was illustrating Oberon-2 complexity for C and C++ programmers by comparing Oberon-2 BNF to their specs same way. Good way to do prelim assessment of protocol/language complexity and whether it's worth the trouble.
It might have been cleaner to specify base64 encoding or length-prefaced payloads (say, 16 bit int preface indicating length in bytes). As it is, you are on your own.
There is no support for a length header or chunking as SSMP is designed with small messages in mind.
This is one of the most annoying things about XMPP (even sending contact photos hits this!), so if replacing XMPP ...
If you're going for telnet compatibility then you'll want to terminate packets in CR+LF, but possibly expect to see only CR or LF from the client (ASCII mode).
Your stream could either be stateful (a message is always sent complete and in order, even if it takes multiple stream packets) or stateless* (different messages might have stream packets consecutively).
It would be more future proof if you started with a message grammar and then defined your protocol on top of that.
Our most common use case is client certificates but there are provisions for alternate auth schemes.