Show HN: Concise Encoding – a friendly data format for human and machine
concise-encoding.org
concise-encoding.org
> Support for the NUL character (u+0000) is implementation defined due to potential issues with null string delimiters in some languages and platforms. It should be avoided in general for safety and portability. Support for NUL must not be assumed.
I believe it's easier to force awkward languages to support nul character than deal with the chaos of not knowing if it's supported or not. They're are other places where they say "this is hard, we'll let implementations to not support it" and it's a ticking time bomb.
Also it doesn't look like you can deduplicate map keys:
> A reference must not be used as a map key.
Which means as much repetition and space waste as in json.
I'm not sure how disallowing a reference for a map key would cause repetition and space wasting. Keys are generally small, so a reference would in many cases end up taking a similar amount of space. Can you talk more about allowing references as keys? I need different points of view NOW, before I finalize v1 of the format! The negative is that it complicates processing, since you could reference a non-keyable value and attempt to use it in a key.
Everybody PLEASE chime in with your thoughts on the format, because I want to make sure I get it as close to right as I can before I finalize V1!
---------- NO. OPTIONAL. PARTS. OF. THE. STANDARD! (And no "but"-s.) ----------
Start there. This might sound flippant and snarky but I promise you that it isn't. Don't ever leave anything to the implementations.
Oh, and build your own compliance tool. So people can just run the results of their own implementations against your tool and see if their code is compliant with your data format or not.
That way lies pain.
I can think of various more compact ways of structuring the same data, but JSON + gzip encodes the naive structure in a very compact way with very little programmer effort.
In general, the design seems to miss a compact representation of an array of homogeneous records, which is a common use case IME.
Edit: Also I agree with GP regarding NUL and other comments about no optional stuff.
I've added array types for arrays of common fixed-length types, but for records it gets complicated since it can only really work if all records are the same size. I could maybe expand array support to custom types for specialized applications that have a lot of record types (provided the custom record type is fixed size). But the primary purpose of the format is a way to allow disparate apps the ability to read each others data in a machine and human friendly way without requiring a bunch of extra pieces. A secondary concern is not trying to be everything for everyone, not complicating the format for too small a gain. But that's the trick, isn't it? (where to draw the line)
But! This is totally a specific use case and you don't need to try make everyone happy. There's GELF and others for structured logs already.
Make everyone’s lives easier and make a spec where everything is mandatory.
I think the best way to deal with this is explicit forking / special-purpose subspec. Where again, in that spec, things are different / extended, but still mandatory. And if consumers fall into that usecase, they can explicitly advertise / consume that version.
date: 4-Oct-2020 4/11/20
time: 13:04:07
url: https://concise-encoding.org
email: kstenerud@gmail.com
ref: @kstenerud
tag: <div color="red">
integer: 123
float: 123.45 1.#NaN -1.#INF
pair: 67x89
percent: 9000%
money: USD$0.50
binary: #{deadbeef} 2#{00101101}
container: [#hashtag "string" %file.ext #(key value)]
...
[1]: https://www.red-lang.org/p/about.htmlThe basic premise is that you can never marry "easy to edit" with "efficient to process". They pull in opposing directions, so the only alternative is to have two 100% compatible formats: one in text and one in binary.
I've chosen the data types carefully to support the most common types that come up in real world situations (and have included custom type support for data not intended for public consumption). The compromises made should allow most people to work without adding extra encoding on top of the encoding (base64, special text parsing and such).
The reference implementation is working for all common cases, and mostly there for array types. Once that's finished, I'll start on the schema design.
2.25 or 1.10
https://github.com/edn-format/edn
Transit is a bridging format to other encoding formats, which is very cool, but a fairly different problem space.
Fressian is pretty close, but is binary only. The main purpose of Concise Encoding is binary-text compatible formats.
bytes = u8[10 ff 38 9a dd 00 4f 4f 91]
You might spend a little more saying what makes your approach unique:
"Use text based formats that are bloated and slow, or use binary formats that humans can't read. Wouldn't it be nice to have the benefits of both, and none of the drawbacks?"
NeXT/Apple property lists have had that feature for decades now, with multiple text and binary serialization formats.[1]
Infra [2] also makes very similar claims "Existing metaformats fit neatly into two categories. They are either textual for human-readability (such as XML and JSON) or binary for compact serialization (such as Thrift and Protocol Buffers). Infra can play the role of either, imbuing each with the desirable properties of the other."
- The binary format is not efficiently packed. For example, Concise Encoding uses 200 of the 256 codes to directly encode integers (-100 to 100) since they are the most common data values in the wild (along with short strings, which also have their own special encoding).
- There's no array type (as in contiguous objects of the same type and size)
- Dates are in seconds, not Gregorian fields, and have no time zones.
- Binary data in the text format is base64 encoded, ensuring that it's 100% unreadable in its raw form.
- Container types are prepended with a length field, making progressive construction impossible.
Infra looks interesting, but it strays from editable text into a kind of hybrid format that requires a specialized editor to read and write. This would only work if such editors became ubiquitous, which is unlikely. I ran down this path for awhile as well, but finally decided against it.
The first formats filled a void. This format joins many players already on the field.
This format fills two voids:
1. The lack of native types in most of the other formats
2. The lack of either editability or processing efficiency, depending on whether it's a binary or text format.
I developed this format because I'm tired of putting encodings on top of encodings (i.e. base64) and other such tricks, just to get my data across to the other side, and I want a general purpose data format that's efficient as well as readable.
It doesn’t seem like this would make transmitting XML or HTML significantly more efficient to transmit compared to just gzipping a string. And if you wanted to use this, you’d have to write the conversion first.
Can someone help me understand the use case for this, ie in which situations this would be super useful?
The overarching purpose of Concise Encoding is to bring forth the tools to make transmitted data accessible to humans, and at the same time efficient to process by machines.
- Binary format is big endian. All modern processors use little endian.
- The time format doesn't have a time zone.
- The binary time format is in UTC, while the text format is in local time, and you have to convert (WTF!)
- The binary time format is HUGE.
- Typed nulls seems a little excessive for very little gain.
- No efficient small integer encoding.
- Containers in the binary format are prepended by the element count, which makes progressive container filling impossible.
- No custom types (except maybe s-exps, which are overkill)
- No metadata
- No URI, UUID types
- no markup container
- no references
Also, which URI standard is part of the spec, what if the URI standard evolves? – It is a complex standard with multiple RFCs and a long history.
What of possible, different length restrictions in language native URI types?
I have no stake in this though, feel free to ignore. I tend to see negatives first.
In XML/XSD 1.1 they gave up on it, and consider any string as valid anyURI
I tried to implement the old XML types. I built a huge regex from the RFC, but it did seem to cover all cases
Does the format validate that a URL is actually a URL?
Why was CBOR not good enough this time? Still not enough string types? Oh, indeed, CBOR doesn't have markup, whatever that means.