Specifying JSON
tbray.org
tbray.org
But then I realized that what I had made up was very similar to type definitions in TypeScript!
Especially if you use `interface` and not `class`, TypeScript gives you a lot of flexilibity in specifying data structures. You get optional values, typed arrays, typed but mixed arrays, references to other named types that you define yourself, and so on. Really, it's all you need and more. The only downside is that you can spec stuff that you can't put in JSON, such as Date objects. As a schema language, that is a bit weird. But just don't spec it like that! :-D
I'll post some examples in a reply when I'm back behind a computer.
For the record, JSON Schemas are quite a bit more expressive than TypeScript interfaces, or even Scala Traits. I actually put together a big list of every JSON Schema constraint that isn't checkable at compile time in TS: https://github.com/bcherny/json-schema-to-typescript#not-exp....
Fair enough on the constraints. One of the design goals I had at the time was that I didn't want stuff like that to be expressible. Basically, it felt over the top to me - you might as well just allow writing predicates in plain JS then. Even more expressive!
My primary goal was a spec language for humans to read and write. Machine validation came second. TypeScript, oddly maybe, appears to have a similar design goal.
I also think it would be great if Typescript offered a "JSON strict interface" that could warn/error if you do something like use a Date instead of a string. (Even better if you could use that type information along with something like decorators for run-time checks before trying to `JSON.stringify`, at least in debug runs.)
The idea was/is to capture the spirit of JSON minimalism, avoiding the insanity of XML-based schema languages.
Currently, some UCSC students are working on javascript and java tools to work with schemas. We should be able to release that work in the next month or so.
And there is (or, better say, was) an "Orderly JSON" language that had nice and compact human-manageable representations, but had essentially died from nobody using it.
JSchema:
{
"name": "@string",
"age": "@int"
}
Pros: is JSON, is simple to read and write. Cons: is (currently) limited to very primitive schemas.JSONSchema:
{
"type": "object",
"properties": {
"name": {"type": "string"},
"age": {"type": "integer", "minimum": 0, "description": "Age in years"}
},
"required": ["name"]
}
Pros: is JSON and reasonably powerful. Seems to be a most common standard out there, with lots of libraries available. Cons: is verbose to the extent it's no fun to read or write by hand. Has some ambiguities and - as the OP article correctly says - spec is unfinished.Orderly JSON:
object {
string name;
integer {0,} age?;
};
Pros: compact and powerful. Cons: not JSON but a custom grammar, no one uses it.There was great work done by OASIS to look into all aspects of meta-modelling and data warehousing, that resulted in specs like MOF and UML2. I'd say that every modern DOM or other object metamodel must be based on these specs (or have some sensible mapping like the one done by Eclipse EMF project for XML schema to Ecore). What's good in MOF is that it easy to map it to programming language metamodels (for example, you can map MOF model to Java Reflection API).
Another option, if you don't need to design it from scratch, is just to use XML schema: JSON is basically a representation of DOM. In Java you can even use the same API to serialize both to JSON and to XML, which means that they both can be described by the same model language. If someone doesn't like XML schema syntax, it's probably possible to load XML Schema XSD into DOM and then serialize it to JSON - voila, you have JSON schema.
Seriously. Everything the others can do, it does better (or, at least as good - not gonna argue that YAML tags are pretty); and most importantly, optionally. (Plus, it's got a few features the others don't)
You don't have to use block text. You don't have to use quotations. You don't have to use tags, or references. A simple usage of YAML is simpler than any other. A complex use of YAML is still less complex than many uses of XML, but has the same information content.
Can you point me at an example of what you're talking about? I can't really evaluate what you're saying with the information I have.
This isn't to say that I would prefer JSON over XML - I most definitely wouldn't - but it is just as abused lately. I hope we will see its demise soon and something coming out of the YAML line (e.g. TOML) getting more widespread.
I only skimmed the Github readme, what makes TOML better than YAML?
That's at the expense of a few things, however. For instance, array can't contain objects of different types; even [ 1, 2.0 ] is not allowed. Also, having a map (aka. a dictionary) in an array is really awkward to write, and there probably are cases that simply can't be expressed.
It has first-class dates, though.
I wrote about the state of settings formats around the time that TOML was born: <http://espadrine.tumblr.com/post/44576550326/a-history-of-se....
And thanks for Dotset - didn't know it exists, but it looks very, very neat.
In most but the most basic configurations, however, I found myself missing some of variables, inheritance, dates. TOML solves the last two, but as you point out it has limitations. YAML solves the first two, though in a bit of an ugly and obscure way. No wonder Salt and Ansible, two heavy users of YAML configs both pass it through templating engines to reach their goals.
For the lack of better alternatives, I usually end up with a Django-style Python module holding the configuration when I want it to be powerful. It still freaks (non-Python) people out though.
Lua is not a half-bad choice if you ask me. It takes some time getting used to, but it's fast to interpret, and at least one project that I know of - Lsyncd - is making an absolutely great layered configuration on top of it.
On top of that schema assisted JSON serialization can be almost as fast as something like protobuf, but you don't have to deal with binary or a lot of codegen on both sides.
PS - I'm not counting the inherent different from simply more characters as important:
<obj> <string>foo</string> <obj>
obj: string: foo
{ "obj": { "string":"foo"} }
That's not important. If you have enough extra characters for that to be important, you have enough data that you should probably do something other than serialize to text.
YAML is missing things such as quote marks so you have to do things such as type inference to determine what an item is, which decreases speed.
I don't know the specifics more than that, but just do serialization benchmarks and you'll see it pretty quickly.
It's possible that's part of what's going on, but... I don't think so. Or, not so cleanly. YAML options aren't choices, they're things you can do if you end up wanting to.
This is a theme in Ruby (and then Rails) programming. There's a bunch of stuff you can use, but you don't have to, and probably won't. There's an effect there, but it's not the same as choice paralysis.
--
I don't see what you see - that YAML was designed with the idea that options are intrinsicially valuable - it said, "Here's a thing to do. What's the simplest way to do this?".
Well-executed simplicity leads to options - think grep. There's a bunch of options when using grep, but really, it just does the one thing... but then there's a bunch of ways to use that one thing.
Again, this is a Ruby thing. There's more than one way to do it because there's more than one situation in which you'll need to. Have a simple array? Brackets. Have a not simple array? Dash.
In either case, it's pretty much the simplest way to do it, but one doesn't work (as well, or at all) for the other situation.
Further, "YAML is not as popular as JSON" is an observational fact, not something we can argue away. It's worth trying to figure out why and learning what we can from it. I've made my proposal. There are other options ("Javascript is a steamroller"). But I'm not really a big fan of the evergreen "other people are just too stupid to see the obvious truth" explanation, which is generally the one that seems to be given to explain why YAML hasn't beaten JSON. That explanation is almost always wrong. (Only almost, though.)
You're saying that the popularity of JSON is directly related to the popularity of implementing JSON, which is due to the simplicity of it's implementation, as compared to YAML. This seems somewhat reasonable, but I don't think it can cover everything.
Sometimes I feel like I'll go to Ruby hell for saying this, but as far as I can tell, XML is actually really good at that part... because it's a markup syntax, not a data syntax, while JSON and YAML are data syntaxes, not markups.
On the other hand, I did a lot with XML in the 00s, and I sure don't want to go back to dealing with that anymore. I'm sure the world has come a long way since SAX was the thing to use for XML, but that experience left me so deeply scarred that I'm willing to just write an ad hoc schema on my frontend and say, "after JSON.parse, go in and coerce these fields to be int / bool / Date, etc". Not great for interop, but what the hey.
Are you thinking something like
{
"business": "Jim's Hat Shop",
"date_visited": #2016-05-03#
}
Or were you thinking {
"business": "Jim's Hat Shop",
"date_visited": "#2016-05-03#"
}The other option ("#date#") is bad, IMO, since it's a magic string. I don't like magic strings because they blow up at unexpected times. Imagine the notice to the users: "please stop using the hashtag #racecar#, we know palindromes are cool, but it is causing our database to crash". I had similar issues with CDATA blocks in XML, actually - escaping strings to afford reasonable parsing is never fun.
On the other hand, the date literal you suggest is not part of javascript. And on that note, why not just use timestamps for dates in json?
With JSON.parse()? That would be a big failure of handling standard JSON if it did that. Do you have an example where JSON.parse() doesn't correctly parse JSON in the manner you suggest? Or really in any manner? I haven't seen it fail yet and if there are known failures it would be good to know about them.
I don't know where the issue really originated which part of transport or serialization, and I honestly don't know how common it really is - I just slapped some type coercions on the object I was putting into the jsonb field (which was necessary anyway, to protect against malicious user input) and called it a day.
So it's probably not JSON.parse at fault here.
EDIT to add: just looked over my notes, and it seems like the problem was likely originating in how I was treating option elements underneath a select in an html form. Somehow the value was getting set to a "10" instead of 10 on the form's model object, and that was getting pushed over the wire (serialized and deserialized) as a string.
The fastest JavaScript implementation we found was is-my-json-valid, but its error messages are nearly useless.
The best error messages (and IMHO conformance to the spec) came from fge's json-schema-validator, which is a Java implementation. The test suite on this impl is impressive. The validation failure messages are excruciatingly long and detailed if you have a large schema, but help a lot with debugging your document generators. But, the performance is awful. If you do any regex patterns, it shells to a Rhino script to to the validations which are predictably catastrophic for performance, making that validator unusable in production environments.
JSON Schema must be very difficult to optimize the way it is structured, which tells me it is probably a badly conceived specification.
JSON in general needs the kind of tooling and library love that XML got back in the day. I've searched in vain forever just for a great query language for JSON documents. I'm convinced that JavaScript+Lodash is the query language for JSON, to which I say meh. And, I'd love to see a first-class JSON streaming API and tooling/libraries (similar to STaX in the good old days of XML). How is anyone supposed to work with huge documents??
We're simply not giving ourselves enough tools, it's all very loosey goosey in JavaScript land.
I can appreciate that an order-independent representation can make schema versioning easier, but what other advantages does this provide?
Python dictionaries hold keys but the keys aren't in any particular order, so allowing them to come out in whatever order maps very naturally to the language.
<myObj>
<Foo>myFoo</Foo>
<Bar>myBar</Bar>
</myObj>
is somehow different from <myObj>
<Bar>myBar</Bar>
<Foo>myFoo</Foo>
</myObj>
is incredibly surprising and frustrating, and XML schema actually supports ignoring the difference but xml schema tools seem to hate using the "all" instead of the "sequence" schema definitions.The worse offender, imho, is Microsoft's DataContractSerializer, whose default behavior is to silently ignore this ordering failure and also silently fail to deserialize any out-of-order elements.
Either fail loudly or be flexible about ordering.
There is no well-adopted analog to xmlschema for json.
<create>...</create>
<delete>...</delete>
Relax NG handles unordered collections well.On a similar note, text protos actually work very well as configuration files and are used by SyntaxNet/Parsey McParseface [1] and Bazel's CROSSTOOL config [2].
[1] https://github.com/tensorflow/models/blob/master/syntaxnet/s...
[2] https://github.com/bazelbuild/bazel/blob/master/tools/cpp/CR...
{"name": str, "age": int, "groups": [str]}
The template has the same 'shape' as the expected JSON, making it very easy for humans to grok.We put all these specs into a file that acted both as documentation and as the actual Python code we used as templates to validate the real JSON. The doc was always accurate ;)
It turned into a full-blown dynamic type checking thing for Python, but it's still most useful for checking JSON APIs at runtime: https://pypi.python.org/pypi/obiwan/1.0.8
If you validate the JSON with an obiwan template, you can avoid having to type-check it as you parse it.
Obiwan syntax would be easy to use in other languages, but we never needed to try.
{
"name": String, // [required]
"age": Number, //int
"groups": [String], // groupId (UUID)
...
}
The structure represents the actual object, and comments allow for additional context/details... it's worked pretty well for documenting APIs that are under development, or intention when working with other devs. data = json.loads(tainted, template=your_spec)
And that will throw an exception if the JSON didn't match the template. From then on its just to safely use the JSON you just loaded, knowing it is well formed and matches the expected shape... :)Libvariant includes a command line tool, varsh ("Variant shell"), that can schema-validate JSON and YAML documents.
There's https://tools.ietf.org/html/rfc7159, co-constraints https://www.ietf.org/internet-drafts/draft-cordell-jcr-co-co..., compact syntax
Unlike XML and UML, there's just less of a market for GUI-like editors for JSON?
Something that would also be nice, if anyone comes up with such a beast would be the ability to edit a multi-record json file... effectively each line is a single JSON, and an LF at the end of each line/record. I've used this structure a few times. Given that JSON will typically encode special characters allowing for single-line representation by default, it's worked really well, especially combined with gzip as a stream for archiving data.
If anyone knows what that program was, would love to find it again.
Tweeted it... It's funny, but I often use twitter to tweet about stuff, just so that I can find them later.
Opening a one-line JSON file of a few MB in Emacs can be painful. The editor struggles to handle the long line and the matching-parenthesis-highlighter runs its stack machine down the whole length looking for the closing delimiter.
There are ways to avoid this, like pretty-printing the JSON across multiple lines, or forcing fundamental-mode when large files are opened, but I still haven't set these up as default.
I believe the question was about good editors for json documents.
i refuse to use XML just like everyone else, but not because it's hard to parse data from.
Although stringly-typed XPath has the same problems as SQL, building complex queries through string concatenation. This is where JSON shines, because you usually just convert to a tree of your language's native data structures in one line and use your languages native facilities to traverse it from there.
object.listOfItems .reduce((item) => item.type === 'car') .map((item) => item.name));
This is way more powerful than xpath.
The advantage of a schema is you can validate your data and move to static typing in only a few lines of code. It's for this reason that I think the best option these days is to just use something like protobufs v3, which has a stable JSON serialization and a sensible schema specification.
But even for objecty data, rather than "documents", XML doesn't have to be clumsy. Checkout the code samples here for C++
The really good thing about XML, that it - unlike JSON - is extensible.
And traversing both JSON and XML (without the weird parts) are no-brainers. Every (or about so) mainstream language has libraries that make either format really transparent to use.
That's really only true about XML, that is not the case with JSON. JSON is a pain to work with in C# for example compared to XML.
I did some toy project in C# and haven't found handling JSON any painful. Maybe I just haven't hit the bad cases, though.
I don't exactly remember what I've used and what I wrote, but I think I had just installed some library (IIRC it was Json.NET), put some annotations on classes, and got my serialization. For deserialization I didn't wanted to write throwaway classes so I just made a JSONPath query and iterated over the result objects, looking at necessary properties. Nothing particularly different from how things are in, say, Python, except maybe for some extra type checking.
[0] https://en.wikipedia.org/wiki/YAML [1] https://en.wikipedia.org/wiki/INI_file
If you do not care too much about exchanging data with other apps then parser/generator supporting javascript style comments (e.g. json-cpp) would work. With each JSON value it can store "commentBefore" and "commentAfter". If you need you can strip these comments any time by parse + write cycle with "collectComments" option inactive.
And I find the whole "run minifiy/strip first" to be a hilarious suggestion from Douglas Crockford. It can be rephrased as "If you want comments in JSON, for a config file say, then don't use JSON."
That, and XML requiring the tag name in the closing tag, are some of the obvious, silly, mistakes. XML in particular. Without the named closing tags, just using, say, <>, would reduce the apparent bloat and annoyance by a very large factor.
'atom-text-editor .insert-mode':
'j k': 'vim-mode:enter-normal-mode'