Uniform eXchange Format (UXF) – plain text human readable typed storage format
github.com
github.com
Instead of:
uxf 1
=Database server:str ports:list connection_max:int enabled:bool
=DateTime when:datetime tz:str
=Owner name:str dob:DateTime
=Hosts name:str
[
(Owner <Tom Preston-Werner> (DateTime 1979-05-27T07:32:00 <-08:00>))
(Database <192.168.1.1> [8000 8001 8002] 5000 yes)
(Hosts
<alpha>
<omega>)
]
Why not: type Database {
server: str,
ports: list,
connection_max: int,
enabled: bool,
}
type DateTime {
when: datetime,
tz: str,
}
type Owner {
name: str,
dob: DateTime,
}
type Hosts {
name: str,
}
[
Owner('Tom Preston-Werner', DateTime('1979-05-27T07:32:00', '-08:00')),
Database('192.168.1.1', [8000, 8001, 8002], 5000, yes),
[
Hosts('alpha'),
Hosts('omega'),
],
]
Evidence shows that people like that kind of format, for "plain text human readable" purposes. They are also used to it. It's only 20% longer, and you can also go with a Haskell-style syntax if you dislike braces.What's the point of a plain-text format that is not human-friendly? Especially for type definitions, do you expect people to write this or do you want them to compile their schema from a different human-readable format into your human readable-format (and why)?
Please cite this evidence. I like your proposed format, just interested in that research.
I agree, but I don't know that why people like that kind of format is settled. I suspect it's because the majority of today's software has to touch the web, and the only programming language built into web browsers happen to consume and produce that kind of format natively.
In other words, I think it's popular, but I think it's popular because it's the path of least resistance for interfacing with the web, which isn't necessarily a priority all the time.
1. It supports data tables with named and typed columns. 2. It supports types in the header that can be referenced elsewhere. 3. It supports lists of stuff as well as types and nesting. 4. It uses a format header to easily declare what format to decode/encode.
Unlike lists of JSON objects, the data can be represented more compactly. I've done something similar when I encode tables as an array of header names, then each row is also an array, where index is used to match the name.
It would be fairly easy to make a binary version of this if you needed more compact representation, and make a lossless conversion between text and binary.
Why would you want a format like this? Many use cases. Every DB wire protocol essentially re-creates something like this, but often poorly. Writing multiple tables that include headers and types as well as data to disk is frequently useful.
The problem with xml schema is XML is really really complex when you add in transforms and namespaces and everything else that XML can include.
The only thing I might suggest is to be able to add meta-data about a specific type, or create specialized type based on type+meta-data (like max length, etc). This could also help with the issue of timestamps (local, second resolution, offset).
Any links to share? I think your feedback is very spot on so curious what you've built.
Wouldn’t it be better to use true for true and false for false?
The only time I can remember seeing yes/no used in a format like this is YAML, and that caused problems[1].
[1]: https://hitchdev.com/strictyaml/why/implicit-typing-removed/
I don't know if I prefer this over treating everything as a string and letting readers/writers decide on their own how to parse things, but it seems like a much more reasonable approach than YAML's.
I'm interested in what you mean when you mention shell scripts; I find it more idiomatic to write:
if [ true ]; then echo test; fi
Over if [ yes ]; then echo test; fi
Because the behavior is more consistent when you try to use it in other constructs: while true; do echo test; done
Versus yes | while read _; do echo test; done
Or maybe while yes | :; do echo test; done
There's also no "no" command. To get that effect, you'd have to confusingly write: yes no
Is there a particular instance of yes/no in shell scripting that you had in mind?Yes!
I see them all the time in various build scripts, especially for Slackware packages / SlackBuild scripts; it's a pretty common convention to use yes/no values for enabling/disabling (respectively) various build options.
OpenBSD's rc.conf(.local) also uses "NO" to indicate that a service/daemon should be disabled entirely; for example, the default httpd_flags=NO in rc.conf entirely disables httpd - unless, of course, you re-enable it later with httpd_flags= in rc.conf.local. "YES" is also sometimes used, e.g. library randomization being enabled by default via library_aslr=YES.
If it sounds stupid when you say it out loud, it _is_ stupid.
(The more locale agnostic ⊤ and ⊥ (https://en.wikipedia.org/wiki/Verum and https://en.wikipedia.org/wiki/Up_tack), IMO are a bit elitist and difficult to type)
I agree that "true" and "false" would be clearer than "yes" and "no", but the fact that "yes" and "no" are English words isn't an issue.
Maybe for communication between trusted parties? Then I would use a binary format.
If human readability was the point, then doing something different than expected is a really bad idea:
* "no" and "yes" as boolean values may save some bytes, but the tradeoff isn't worth it (and if filesize matters, use a binary format to begin with). * Using angle braces except of double quotes to fence strings makes the format look noisy and means you have to remember two kinds of escapes if you want to use < and > in the value. * The format isn't object oriented in any way. You can simulate that by putting maps into maps, of course, but no one will have fun reading or writing that in a text editor.
Type information is for parsers, not humans. JSON this this right, Protobuf does this right. UXF is just a compromise combining (only) the disadvantages of the two.
UXF is self contained, that's great, but in 99.9% of the cases where you need a DX format, sender and receiver already know the schema, so that definition block just adds bloat.
You can happily mix lists, maps and tables of primitive or compound types. And since stuff is typed instead of named, order matters and you end up addressing everything through positional parameters. That's going to be fun when using a text editor to write down something like a list of GPS coordinates (you are likely to confuse latitude and longitude).
I feel the spec warranted more discussion about why strings <look like this> instead of the way more common "like this", i.e. why angle brackets are used to quote strings. Probably to make it easier to embed quotes, but I'm not sure. It was rather surprising at least, although I guess you get used to it if you read a lot of raw files.
https://github.com/tlocke/zish
Any comments / criticisms gratefully received.
For example, imagine I have a compact append only map format and I want to represent it like this (where last tuple wins when you have a duplicate key, but earlier history is preserved)
{
score: 0,
score: 1
}Chesterson's Fence (https://en.wikipedia.org/wiki/G._K._Chesterton#Chesterton's_... is a very powerful design principle. They chose to put those elements into ISO 8601 for principled reasons: they come from pain. They embody responses to mistakes that I've made, and thousands of other engineers before me. Unless we fully understand the reason they were included, don't arbitrarily to do "I haven't used it, so it must be useless."
Other than that, it looks like a clean spec, but I'm not personally convinced that it has enough incremental value over JSON or YAML to replace them in the human-readable exchange format space. It can be a little more concise, but if I'm making something for humans, clarity (typically) has more value than conciseness. Are there other compelling values that I'm missing?
Anyone saying “UTC” is wrong. Unambiguously wrong if offsets are supported, and in foolish contexts like this where offsets are not supported, still wrong due to common sense and custom.
If there is no offset, there is no offset. It’s what is commonly called a naive or plain datetime. How it should be interpreted is explicitly undefined if offset-capable, and implicitly undefined by strong custom if not offset-capable; but it will generally mean in the local time zone, whatever that is—and it could be relative to a particular machine or a particular user. This is often suitable for social use, but completely unsuitable for machine history-recording use.
So: the question is rhetorical, unanswerable, thereby demonstrating why nmz’s position is unreasonable.
(Actually, only probably unreasonable because nmz’s wording wording with its “X” and “y” is not clear and may be using the term “timezone” subtly—the trouble is it’s used to mean three different things: firstly and most properly, a name for a set of rules about which time offsets to use when, e.g. “Australian Eastern Time” or “Australia/Melbourne” as it’s called in the IANA Time Zone Database, which roughly means AEST (+10:00) for half the year and AEDT (+11:00) for the other half, but conveys the rules as they have been through time; secondly, a somewhat less correct colloquial usage, a named time offset, e.g. “AEDT” or “Australian Eastern Daylight Saving Time” for +11:00; and thirdly, fairly clearly into the realm of misuse but still very common, a time offset like “+11:00”. If nmz was using the term “timezone” more precisely to mean one of the named concepts and expressly not an offset, then yeah, times written that way do require memorising a whole database, whereas offsets are straightforward to calculate, though it’s definitely harder having to do two calculations than the just one if it starts at UTC.)
—⁂—
For the rest of your statement: for times not tied to a particular location or time offset, you should always use UTC in the form of the offset Z, in ISO 8601/RFC 3339 terms, since specifying any other offset indicates that it means something. (Note that RFC 3339 tried to have -00:00 be the neutral offset and Z and +00:00 meaningful, but that is acknowledged to have failed, and so https://www.ietf.org/archive/id/draft-ietf-sedate-datetime-e... is updating it to match actual usage.) But for things that involve humans and are anchored to a particular time zone, using UTC and not storing a time zone is wrong: you should store the relevant time zone and (fallback) offset so that if the time zone definition changes (as they do, sometimes with less than a few days’ notice), future times can be corrected, which they can’t be if you anchored them to UTC. So: things like system logs, use UTC; online conferences, use UTC; location-bound conferences, use that location’s time zone; general user calendars, use the user’s time zone; calendars for companies that straddle time zones (or people that work across time zones): deliberately choose a time zone or offset to anchor things to (sometimes at the level of individual events), especially for the sake of recurring event periods if you use a time zone with DST.
While a datetime without a timezone isn't terribly useful, distinguishing it from a datetime and timezone designation pair is the only correct type system.