Z85: Format for representing binary data as printable text
rfc.zeromq.org
rfc.zeromq.org
The five characters SHALL each be converted into a value 0 to 84, and accumulated by multiplication by 85, from most to least significant.
There is of course the example and the reference implementation and everyone knows how this kind of encoding works, so you can certainly figure out how you are supposed to implement it.
Anyway, it all seems a bit https://xkcd.com/927/
Because there is no padding rule given, it's impossible to use the standard to encode most binary strings.
For example, how would you encode/decode these?
0x0000000000000000
0x00000000000000
0x000000000000
0x0000000000
0x00000000
0x000000If people would start using this encoding, different users would adopt different solutions padding, length prefixes etc. and it becomes mess.
It's not "impossible to encode" that content just because you need to decide how to represent it. It's not "a mess" if some people use fixed-length strings and some use length-prefixed strings. It's just reality for any encoding scheme - you build layers around it, according to how you want to use it.
The same way the context of your software determines whether this binary string is supposed to represent a float, a name, a hash, or a pixel-art masterpiece, it will also determine the appropriate serialization.
I myself typically binary compress then base64 encode.
I use 85 in json just fine without escaping overhead. Never tried for URLs or XML.
The article explains how Base64 is “problematic because it has more than a dozen variants”. So instead of picking one, let’s invent a new standard.
> Fortunately, the charging one has been solved now that we've all standardized on mini-USB. Or is it micro-USB? Shit.
is hilariously even more true.
Example: 0 to 4294967295 = 4 bytes
4294967296 to 4311744511 = 3 bytes (subtract 4294967296 first)
4311744512 to 4311810047 = 2 bytes (subtract 4311744512 first)
4311810048 to 4311810303 = 1 byte (subtract 4311810048 first)
Then you're still left with 125242821 remaining numbers, over 26 bits.
Another is to encode 4k+i bytes as 5k+1+i base-85 characters, for 0<i<4. That way the encoding length immediately determines the input length. And there's again plenty of space since 85^{i+1} > 256^i for i < 4.
This leaves encoding lengths of 5k+1 unused. These could be used to support arbitrary bit lengths, i.e. for encoding an input of 4(k-1) bytes + i bits, with 0<i<32. Set the final byte to i, and let the preceding 5 base-85 characters encode the i bits with 0 padding.
Why oh why??!
If it were little endian, you could probably skip the "must be multiple of 5 chars/4 bytes" requirement, not to mention that 99.9999% of processors out there are running in little-endian mode.
There is nothing "envious" about network byte order.
Assuming n is an integer:
* 5n bytes received = 4n bytes data
* 5n+1 bytes received is [invalid]
* 5n+2 bytes received = 4n+1 bytes data
* 5n+3 bytes received = 4n+2 bytes data
* 5n+4 bytes received = 4n+3 bytes data
This is like modified Base64, which doesn't need any padding.This comes down to whether there should be 5 valid encodings ("10000", "1000", "100", "10", "1") of a single 0x01 byte, or one. The variable length encoding of integers in Protocol Buffers has the same malleability problem
It's also not clear to me why you say 6 char input is invalid.
Those are the same if you're treating the binary data as a stream of 32 bit numbers, but not if it's a stream of an arbitrary number of octets.
Your parent is suggesting that if after chunking the input into 5s, your last chunk is "10" you would treat that as 0x01, "100" as 0x00,0x01, "1000" as 0x00,0x00,0x01 and only "10000" as 0x00,0x00,0x00,0x01. That's not four encodings of the same value at all.
Treating "1" (or any single leftover character) as invalid in such a scheme makes sense because a single character can only encode 85 values, from 0x00 to 0x54.
Then I want the receiver to understand that only the first 6 bytes of their decoded results are part of the transmission -- how do I do that?
Base64 has a special character ('=') that is used for encoding padding, but this method doesn't seem to have that. The spec says "it is up to the application to ensure that frames and strings are padded if necessary", which suggests they've scoped this problem out.
I suppose I can always build a little "packet" that starts with the payload length, so that the receiver can infer the existence of padding if there is additional data beyond the advertised payload length, but now the receiver and I need to agree on that protocol.
Unfortunately for Z85, they made the highly questionable decision to use big-endian, which means it can't take base64's route. You could probably define an incomplete group at the end to be right-aligned or similar, but you may as well be sensible and just go little-endian.
Not including padding will seriously hinder this. You want that UX for people to be able to use it widely and just work. The post talks about wanting one compatible standard, but leaving out how to deal with padding means you know have incompatible ways of doing so. Plus many junior devs won’t even understand the need of padding, and will be extra confused when it doesn’t “just work.”
I do appreciate the many other bits it tries to solve, however.
Also this spec:
> It is up to the application to ensure that frames and strings are padded if necessary.
So they didn't even specify the standard way to treat byte strings with the length not divisible by 4.
GPL is just a particular set of terms.
GPL never meant no copyright or no terms. In fact it sets very strict terms, just not very many, not very complicated, and not the usual ones.
You never knew the basic premis or theory of how the gpl works?
> This Specification is free software;
even though This Specification is not software.
And then you're left with the disturbing need to send the text of the GPL off to corporate lawyers to determine whether there are any bombs in the GPL when applied to a specification instead of a piece of software? Will implementing the specification contaminate our sources? Must we include a GPL 3 notification in our legal notices?
To "modify" a work means to copy from or adapt all or
part of the work
in a fashion requiring copyright permission,
And the corporate lawyers will respond (as they always do) that the text of the GPL wasn't written by a lawyer and nobody really knows what it means because of numerous drafting errors.Using the GPL here senselessly and needlessly causes stress, and is a more than ample reason in itself not to use the standard.
The one good thing is that they used an MIT license on the reference implementation, so we won't have to use clean room protocols to write our own implementation.
This opens up a new question: so this encoding will never be allowed in proprietary software?