Concise Encoding: A secure data format for a modern world
concise-encoding.org
concise-encoding.org
- Versioning.
- Time zone identifier instead of just a fixed offset (which are ambiguous for future events).
- Native encoding of binary values.
- Graph notation with support for labels.
- Comments!
- Trying to escape lookalike characters, even though I think that's a lost cause.
Things I'm not so keen about:
- NUL character in strings being platform and settings dependent.
- Line break not being forced to a consistent value.
- The most complicated number encoding scheme I've ever seen (e.g. 0xa,3fb8p+42).
- Entity references are a footgun for anyone writing a depth-first or breadth-first algorithm.
- Arrays-vs-list feels like it doesn't belong in encoding formats.
Hexfloat can be really useful when you need precise/exact floating-point constants for numerical methods. Without them, you end up having to do more-complicated hacks to preserve exact constant values when code gets compiled, or you have to live with compilers (sometimes) subtly altering constants.
I wish more languages supported hexfloats.
You can also use them as floating-point literals in your code.
Time zone identifiers are not future-proof. As a concrete example imagine a fictitious country X with the largest city (not necessarily the capital) being Foo and a single time zone throughout the country. A natural time zone identifier would be something like `Region/Foo`. If a part of X switches to another time zone and Bar is the largest city in that region for whatever reason, there would be another time zone identifier `Region/Bar` assigned. You can't automatically determine which `Region/Foo` should be converted to `Region/Bar`. It is up to you to pick a strategy in such cases, but time zone identifiers themselves do not solve the ambiguity problem.
Besides from this false sense of security, time zone identifiers were not meant to be for general interchange and there are several understandable but problematic assignments. Baking them into a data structure format doesn't sound a good idea.
Fun fact: time zones are named after cities for robustness:
Country names are not used in this scheme, primarily because they would not be robust, owing to frequent political and boundary changes. The names of large cities tend to be more permanent. Usually the most populous city in a region is chosen to represent the entire time zone, although another city may be selected if it is more widely known, and another location, including a location other than a city, may be used if it results in a less ambiguous name.
https://en.wikipedia.org/wiki/Tz_database#LocationA self defining and self describing set of measurements would help here but obviously bloats formats similar to how XML schemas become overly verbose and pedantic without adding clarity to the target audience.
So I agree that baking in a time zone is likely a bad idea but I have doubts we can do much better by trying to use another more absolute time coordinate system independent upon human civilization conventions.
I think of a time zone as drawing a large box on a map, for any given instant in time. Future events either happen 'only in time' or they happen at a time and place. A time zone isn't a great representation of a place! I'll be much happier storing UTC and coordinates, then turning that into a timezone for display when I have to.
A time zone is an approximate (and weird) spacial coordinate system which causes nonlinearities in the representation of time, and unless you want both of those properties it comes with baggage.
That approximate and weird system is how humans tend to think about future events though. You're right that often it is not necessary, but for many usecases (e.g. when a TV show will be first broadcast in the future) they are the most robust model we have.
My assertion is that a very strong argument should be won before they are used for any other purpose.
The problem was that the calendar advertising the recurring event was set to PST. And daylight savings had changed in California (PST -> PDT or the other way around). The result was that the calendar event shifted by 1 hour local time for me and every other attendee who subscribed to the calendar event.
This situation sucks, and its really confusing to everyone. But I can't think of a way around this whole issue:
- If recurring calendar events weren't set to a timezone (and thus, just worked off GMT), then all recurring calendar events would drift forwards or backwards when daylight savings changes happen
- If recurring calendar events are set to some local time zone, then things like this happen - which really confuses everyone involved.
I mean, sure - in an ideal world we'd get rid of time zones. But until then, it seems like we're stuck with this problem.
(Though at least daylight savings time is slowly being phased out in some countries.)
Calendar apps are exactly where the full and monstrous complexity of time zones emerges, and I don't envy anyone who works on one.
My case is that the problem doesn't have to happen in the other direction and this unforced error is made frequently. There are a huge class of recurring events where drifting by an hour on the local clock won't make a difference, but either skipping something or doing it twice is bad.
What you're describing strikes me as a bug in the absolute sense: nothing should be displaying PST during a time frame when that time zone isn't in use.
Again, don't sign up to have these problems if you can possibly avoid it.
There is the extra problem that the standard explanatory names for timezones that appear in the CLDR are very confusing, eg British time is referred to as GMT even though that is wrong for more than half the year in the summer.
†arbitrary, yes, but real
I'm saying something stronger than it often isn't necessary, I'm saying that storing time zones as part of a time is almost always the wrong thing to do.
For example: if your show is broadcast first in Eastern then in Central time, that's one thing. What if it shows in Canada the next day? All of a sudden you really wish you had modeled space separately from time.
Even when some event is time-zone-gated, this will correspond to one exact moment in time, and for e.g. server provisioning that instant is what matters. Including a time zone can only lead to missing that instant, it can never help you find it.
I even added a write-up about time to the spec: https://github.com/kstenerud/concise-encoding/blob/master/ce...
[1] Back when I designed my own serialization format I criticized TOML's decision to add date and time types for the same reason: https://github.com/lifthrasiir/cson#no-additional-types
I know this sounds like an "if everyone would just ..." kind of excuse, but poor time implementations really are a big problem in the industry, and I'm hoping to at least provide the tools for knowledgeable people to do it right.
Line break is another sticky issue. Non-technical Windows users will inevitably produce documents that use CRLF, so if it rejects the file, that's a bad user experience. What's the best trade-off here? I'm not really sure.
The number encoding thing is being discussed in other comments so I'll leave it be. I had allowed comma because that's how a huge chunk of the world represents the decimal separator. Once again, this is about user experience with the non-technical users. I'm REEEEEEEALLY on the fence with this one.
Entity references are dangerous, yes, but also powerful. The point of them in the format is to solve the recursive reference problem, because you just can't do that otherwise, and these structures do exist in the world. It's another case of an imperfect solution for an imperfect world. Bear in mind I absolutely do NOT want this format to become some Turing complete language. This is just the minimal feature set I could think of to represent real world data.
Arrays-vs-list is another one of those compromises. Encoding an actual array of fixed types into a list would be slow and bulky, leading people to just encode them as a chunk of bytes like they currently do in JSON and other formats. As an imperfect solution to an imperfect world, I want to at least let people preserve the semantic meaning of what they're sending since they're going to use array encodings regardless of what the format supports.
> NUL is a tricky beast
One alternative I found interesting is Modified UTF-8 (https://en.wikipedia.org/wiki/UTF-8#Modified_UTF-8), where NUL is encoded as 0xC0 0x80. It's already in use in Java, apparently. Disclaimer: I've never heard about it before, and I don't know how well supported it is.
> Line breaks and commas
I'm from a country that uses comma as decimal separator, and I still prefer dots when programming. I dread ambiguous numbers like "1,001". I'm confident this is true for almost all technical people. And I really don't see non-technical users editing this kind of file.
If they can be trusted to read it, and to modify it without introducing syntax errors, they can be trusted to use dots and an editor that shows LF as linebreaks (i.e. anything but `notepad`).
My choice would be dots only for decimal separator, accept LF and CR LF on reading, but prefer LF when writing.
> Entity references
You can always define manual ids:
c1
{
"alice"={friends=["bob" "charlie"]}
"bob"={friends=["alice"]}
"charlie"={friends=["alice"]}
}
Of all data types that could have native encodings (colors, IP addresses, lat/lon coordinates, enums, tags, markdown, hashes, file permissions, OIDs, DOIs, ISBNs, etc), I think arbitrary object graphs bring too many downsides for a serialization format.You don't want reader CVE's because of Billion Laughs, and you don't want to force every programmer writing a simple traversal algorithm to correctly handle cycles.
Damn, that's a clever idea!
> I dread ambiguous numbers like "1,001".
Good enough for me. Commas will be removed.
> accept LF and CR LF on reading, but prefer LF when writing.
That's the current behavior.
> and you don't want to force every programmer writing a simple traversal algorithm to correctly handle cycles.
Ugh... I really really REALLY want you to be wrong on this :(
> That's the current behavior.
I'd suggest changing the SHOULD to MUST here, and remove the "foreign or unknown system" part:
but encoders SHOULD output LF when the destination is a foreign or unknown system.
It's ok if a CR LF sneaked in because a user edited a file manually, but encoders should be more predictable.>> and you don't want to force every programmer writing a simple traversal algorithm to correctly handle cycles.
> Ugh... I really really REALLY want you to be wrong on this :(
I have some good news then.
I just checked, and most of my JSON traversals are for things that you already take care of, like binary arrays and handling cycles (huh, talk about irony).
And the billion laughs problem was mostly because XML entities are more like macros, and expanded in place. As long as the reader doesn't try to convert the document to JSON, or naively print the object graph, it should be ok.
I think it might be ok to keep references.
And again, cheers for the encoding specification! It's really cool, and I hope it catches on.
Personally, I disagree with how your format handles all those other issues too (except for the numbers), but well, if you think you are correct, go try it. If it works, it works, and my disagreement may easily be misguided. Anyway, I disagree because:
For the line breaks, the internet has a way of trying to "fix" them and completely breaking the line-information of the original document. It would be ok if the format wasn't blank-space dependent, but it is, so changing the lines breaks the data. Anyway, that is becoming a lesser problem with time, so maybe for a new format it's fine.
Entity references on formats that are not focused on them are surprising. That means a lot of software will break once they get one, and tradition says they will do that in a way that compromises computer security. I would either change the format so that references are almost always used or remove them. If an application needs references, it can always tag the entities with an id and put the references there by itself.
The same applies for arrays, in a lesser degree. They will be surprising, but they are also easier to handle. But they are also much less necessary, since lists can always replace them. I'm really not sure if they are a net negative or positive.
Rest assured that this project IS an ongoing concern, and is the foundation of many more technologies I intend to bring to bear over the coming years.
Are you open to feedback in the spirit of collaboration? I've considered opening an issue or two on the repo, but it always feels a bit presumptuous to do that when one person is the author of the entire project.
I like the way you think about these things and I find CE quite promising.
One of the main reasons why I haven't released version 1 yet is because it's been so much of a one-man-show that I'm under no illusions of being immune to tunnel vision.
In the immortal words of Johnny 5: I need input!
Too many question marks in my head to be honest. The more I think about it, I believe code still may be the the best way to describe really complex data.
What is interesting to me is the holistic view of all the data types we most commonly use--but then I don't really know how this differs from Thrift/Avro/etc--and then further, why haven't we as a community moved to one of those?
{"name":"name1", "phone":"+769989823"}
{"name":"name2", "phone":"+769244563"}
{"name":"name3", "phone":"+769989295"}
...
There should be a way to do the same thing with concise.
no they are not lol