Preserves: An Expressive Data Language
preserves.dev
preserves.dev
Not a fan of annotation (that can be used as comment syntax) having # followed by a space character have a different behaviour feels strange.
Until a project has a lot of traction (think docker, react, django not uv or jq) it's very safe to assume that every visitor to your page doesn't understand the background.
Except they do in Python. It is extremely useful, surprisingly often.
(ETA: What are you quoting there? I don't think that text appears on the Preserves site) (ETA2: Ah, it's the tutorial. Cool)
Plain dict maintains insertion order but equality checks only check that the key/value pairs are the same. [2] [3]
[1] https://docs.python.org/3/library/collections.html#:~:text=e...
[2] https://docs.python.org/3/library/stdtypes.html#:~:text=dict...
[3] https://docs.python.org/3/library/stdtypes.html#:~:text=dict...
Interesting on how on one hand the size of SignedInteger is unlimited, but on the other hand there is a ByteString. A ByteString could also have been represented by as a sequence of SignedInteger. I also wonder if it would not better to have a Unicode character as an atomic unit and represent a string as a sequence of Unicode characters.
This makes me wonder whether this is a high-level data model or yet another data representation.
<tag v1 v2 v3>
If you put a single dictionary-valued "field" in a record, you get a variation with named fields <tag {
field1: value1
field2: value2
field3: value3
}>
Records have positional "fields" because of the Scheme heritage of the design.--
Re bytestring -- yes there are some concessions to real machines/languages in there that aren't absolutely required. Other examples include booleans and strings, which could have been <true> and <false> and <string [65 66 67]> etc respectively.
There's a little more on this topic in footnote 2 on the "conventions" page: https://preserves.dev/conventions.html#fn:why-dictionaries
Are there any good examples of nontrivial schemas etc?
The syntax isn't the most interesting part though; the thing that distinguishes it from most other data languages out there is that it has semantics (= a rigorous definition of when values are equal and when they aren't). So you can use Preserves semantics with JSON syntax (a subset of Preserves' text syntax) as one way of getting actually-meaningful JSON.
Plus, comments (and other annotations) ;-)
As someone familiar with Protobuf, comparing Preserves vs Protobuf text format, here's my quick comparison between the two after reading through the tutorial:
- Preserves' Symbol is very close to Protobuf enums. But Symbol can contain characters like dash
- There doesn't seem to be an equivalent of Preserves' Record in Protobuf, but the tutorial's example of using <Unknown ...> To denote a missing <Date ...> can be simulated using the `oneof` field in Protobuf.
- Having to write #t/#f in Preserves is unfortunate. I guess this is the result of schemaless serialization language and potential parsing ambiguity with a Symbol?
- Protobuf have a way to annotate the schema and reuse at runtime, very similar to Preserves' annotations.
It looks like Preserves just uses version numbers in its schemas. On the other hand, you can read the data without a schema, similar to JSON.
The schema language is extensible/evolvable in that pattern matching ignores extra entries in a sequence and extra key/value pairs in a dictionary. So you could have a "version 1" of a schema with
Person = <person @name String> .
and a "version 2" with Person = @v2 <person @name String @address Address>
/ @v1 <person @name String> .
Then, Person.v2 from "version 2" would be parseable by Person from "version 1", and Person from "version 1" would parse using "version 2" as a Person.v1.The schema language is in production but the design is still a work in progress and I expect more changes before a 1.0 release of the schema language.
(The schema language is completely separate from the preserves data model, by the way -- one could imagine other schema languages being used instead/as well)
Protobufs have an extra level of indirection built in: code refers to fields using names, but numbers are sent on the wire. Without convenient access to field numbers, they can’t as easily be hard-coded. This also strongly encourages using the schema file for most tasks. With protobufs (or similar), any user-friendly editor will need a schema to make sense of the data.
JSON-like systems and protobufs have opposite design goals: encouraging versus discouraging schemaless data access.
Person = <person @name String @address Address>
as above, or Person = <person {
@name "name": String
@address "address": Address
}>
or Person = {
@name 1: String
@address 2: Address
}
etc. all produce the same host-language record, e.g. in TypeScript export type Person = {
name: String,
address: Address,
};After some brief reading of docs, I'm trying to write one sentence explanations. Maybe this will be helpful to you
What
Preserves is a specification and set of libraries in popular languages that lets you reliably exchange data between XML, JSON and EDN.
Who
Preserves is built for (data engineers|data framework writers) to reliably interchange data.
Why
Formats like JSON in particular are imprecise. Preserves forces you to deal with these vagaries up front
What else?
With P-Expressions you can search a preserve compliant datasource much like you would query JSON with JQ
Who Not? Who shouldn't use this
This will not help a data analyst exchange data between CSV and Excel
The name nevertheless feels awkward to me, also in spoken conversation. A made-up word like maybe “Pres” or “Edal” (from “expressive data language”) would work better IMO.
The reason JSON is lower friction than XML for data representation is that you get basic data representations (numbers, strings, arrays, maps) for free in a natural native syntax that happens to parallel multiple programming languages.
XML, in contrast, is a meta-language that allows schema to express different data representations. You've got to use attributes and elements to represent data and data types. XSD is a common datatype schema, but it's quite verbose, and data serialization looks very different from what it looks like in a programming language representation.
Preserves looks like a superset of JSON. It includes additional data representation concepts through syntax extensions, but the idea is the same.
What I don't see is a standard way to map record types (like "irl" in the tutorial) to a unique identifier like an URI/IRI, or something like a CURIE. That kind of feature would allow Preserves to better describe standardized record types.
Decimals I'm on the fence about. Some discussion here: https://gitlab.com/preserves/preserves/-/issues/10