SGML derivatives make sense for markup, not structured data representation.
wait, what?
Both YAML and JSON map to data structures in your program; scalars, sequences and mappings.
XML on the other hand maps to.. a DOM! You've still got to normalize that into what your program expects. SAX is not any better.
Everytime I land in a project using XML I know I'll spend orders of magnitude more time dealing with data than if the project had used YAML or JSON.
> XML on the other hand maps to.. a DOM! You've still got to normalize that into what your program expects. SAX is not any better.
Of course, your program probably operates on more than just scalars, sequences & maps, which means you still need to convert your JSON structures into application data structures — validating along the way, too.
If you have distinct application structures that means you're reinventing ad-hoc scalars, sequences and mappings for the sake of encapsulation. You gain absolutely nothing yet you now have mountains of rigid code to maintain and you can't use generic data functions anymore - more ad-hoc stuff to write!
Even working in statically typed languages you can get away with HashMap<String, Object> or IDictionary<string, object> and call it a day. Very few things require the performance of custom types and its a complete myth that types makes you safer - you have so much more code that bugs are more likely to crawl in.
No, it means that one has types. What does '{"length": 3}' mean? 3 items? 3 inches? 3 kilometres? For that matter, is '{"length": "abc"}' incorrect?
> Very few things require the performance of custom types and its a complete myth that types makes you safer - you have so much more code that bugs are more likely to crawl in.
Having caught many, many bugs from strong types, I feel comfortable saying that you're incorrect here. Yes, there are even more sources of bugs than just type errors. There are also more bugs than memory-management errors. But just as manual management of memory very rarely makes sense, manual management of types very rarely makes sense.
If you don't have a solid sanitisation/marshalling/untainting/whatever-you-call-it layer, you're going to have security and other correctness errors.
{length: 3} has a very specific type. Shoving that into a struct Foo { double length; } doesn't make it any safer. A simple validator and properly naming fields to mention units makes you safer than types.
> For that matter, is '{"length": "abc"}' incorrect?
That's a non-issue with even the simplest of validators.
> Having caught many, many bugs from strong types, I feel comfortable saying that you're incorrect here.
Yet you still need a test suite to handle the huge holes left by the type-system. These tests will catch all bugs the strong types would along the way. Unless you're having a Hindley-Milner type system, chances are you're working in C++, Java or C# and bugs caught by these type-systems are very superficial.
> If you don't have a solid sanitisation/marshalling/untainting/whatever-you-call-it layer, you're going to have security and other correctness errors.
Totally agree, but you don't need types or XML for that.
Having worked as much with statically typed languages and dynamic ones, I'll chose the later any day unless I have very strict performance/memory requirements which is not happening for 99.9% of projects.
Types might catch the easy and superficial bugs, but my ability to reason prevents the hard ones. Having orders of magnitude less code to think about is exactly what I want to achieve it.
> That's a non-issue with even the simplest of validators.
That's the point: the user-marshalled data must pass through the validator, and once it's passed through, it's no longer the same type.
Perl figured this out with its taint mode 20 years ago. We shouldn't be still exposing ourselves to these sorts of bugs.
> Unless you're having a Hindley-Milner type system, chances are you're working in C++, Java or C# and bugs caught by these type-systems are very superficial.
If a better tool is available, why use a lesser?
> Totally agree, but you don't need types or XML for that.
I'm not arguing for XML. XML is almost as horrendous an embarrassment for the profession of software development as is JavaScript. If someone is looking for a data-interchange format and settles on XML, something is deeply wrong with either him or his situation. XML's just terrible.
But passing around JSON objects within one's code is insane. JSON is a (decent) data-transfer format, an completely unsuitable for computation.
> Types might catch the easy and superficial bugs, but my ability to reason prevents the hard ones. Having orders of magnitude less code to think about is exactly what I want to achieve it.
I completely agree, which is why I don't use XML, JSON, JavaScript, C or Java. It's why I use a marshalling layer: I only have to think about user-marshalled data in that layer; I get to only think about my problem set everywhere else.
How is that different from loading a JSON into custom data structures? You run the exact same validation step, have the exact same data for input, but your result is now unusable by the entirety of your generic data functions. You'd be better off with a hash-map than a class.
> If a better tool is available, why use a lesser?
Because no tool is better all across the board. Otherwise we'd have the holy grail of programming languages right there. However, I'll favor productivity and the ability to reason about code over performance or some weak validation that I didn't make a typo or sent a string where a boolean was expected.
Besides, I have already spent weeks trying to debug type errors in C++; its ridiculously easy to break the type system. I've also had countless cast exceptions in both Java and C#; how are types helping there? If anything they give you the illusion of correctness.
> But passing around JSON objects within one's code is insane.
That's not at all what you do. When you do JSON.parse the result is already using the language's data structures; there is no concept of "JSON objects". Even in C#/Java you can map these to dictionaries or use reflection to directly map these to ad-hoc structures/classes.
> I completely agree, which is why I don't use XML, JSON, JavaScript, C or Java.
Its rarely up to you. There aren't many endpoints supporting something other than JSON or XML. Anything running in the browser is automatically running JavaScript; anything embedded will go for C.
Once the data is past the validator, its no longer user-marshalled data. Its just data. Wrapping it around ad-hoc types doesn't help your problem domain, it just adds yet another level of indirection you could easily do without.
[edit] I should also add that I strongly favor YAML over JSON for two simple reasons; the format is both versioned and extensible. That fixes the problem in JSON of using string conventions for fancier types like dates. There I can load the data directly in exactly the shape I want it to.
That's what I'm arguing for doing. I thought that you were arguing for using JSON data structures rather than custom data structures.
> You run the exact same validation step, have the exact same data for input, but your result is now unusable by the entirety of your generic data functions.
What do you mean by 'generic data functions'? I can do whatever I like with whatever parts of my language I wish.
> Besides, I have already spent weeks trying to debug type errors in C++; its ridiculously easy to break the type system. I've also had countless cast exceptions in both Java and C#; how are types helping there? If anything they give you the illusion of correctness.
They're not: C++, Java & C# are all examples of lesser tools. That's why I don't use them.
> When you do JSON.parse the result is already using the language's data structures
Which is what's insane. I don't want a dictionary; I want an instance.
> Once the data is past the validator, its no longer user-marshalled data. Its just data.
And we discovered decades ago that it's useful to talk about categories of data (types) and about data that carries along with it behaviour (classes). Why limit yourself to whacking things with a stick to collect pinecones when you can use a starship to explore the galaxy?
> Wrapping it around ad-hoc types doesn't help your problem domain, it just adds yet another level of indirection you could easily do without.
Programming is all about building intelligent abstractions which enable you to reason at a higher level. I don't want to think about the third item of an array which is the value of a property name 'foo'; I want to think about interest rates, organisational rules and categories.
I like YAML too; it's a decent format. Not perfect, but decent. Bit big though. I'd prefer something lighter weight like S-expressions though: just enough structure to make it easy to turn into native objects.
I very rarely go for custom data structures :) I won't destroy data by giving it an ad-hoc type, and unless I'm working in a statically typed language I won't have a "JSONValue" type to deal with in the first place; I'll get the same value from JSON.parse as if I wrote "const defaultCfg = {host: 80};".
But now I can do something like "const cfg = merge(defaultCfg, validateConfig(JSON.parse(cfgFile)));" where validateConfig takes a hash map and returns a hash map (or something behaving like a hash map).
[edit] by destroy I mean you lose the ability to work on unknown data, which is a requirement to build robust distributed systems. When you marshall into a custom type and lose 3 fields, the endpoint is not going to be happy.
> What do you mean by 'generic data functions'?
Functions like map, filter, reduce, take, drop, zip, split, join, get, nth, head, tail, and hundreds more operating on scalars, sequences and mappings.
> Which is what's insane. I don't want a dictionary; I want an instance.
A dictionary is an instance. Your custom class with getters/setters is nothing more than an ad-hoc dictionary.
The only time I'll go for a struct is when I get performance constraints and need to pack/align data and have continuous memory to batch operations and not trash the cache.
I don't want a custom type where I have to write boilerplate around for every single operation when I have hundreds available ready to be composed.
> And we discovered decades ago that it's useful to talk about categories of data (types) and about data that carries along with it behaviour (classes).
Data by definition has no behaviour, its plain information. You're talking about state and there I agree, I want my state close to its behaviour because they're tightly related. But I don't want the _declarations_ of state and behaviour to be coupled either. Haskell and Clojure both solve this problem nicely with type classes and protocols respectively.
You can make your types behave like sequences and mappings, but that's just extra boilerplate when talking about data and not state.
> Programming is all about building intelligent abstractions which enable you to reason at a higher level.
I agree, but you only need a handful of abstractions, most of the time what you're doing by wrapping data in a type is an indirection not an abstraction.
S-exprs by themselves are just data, they're the very same as JSON (not versioned, not extensible). YAML supports tags making it easy to plug your custom types directly into its reader.
I'd say that a dictionary is an ad-hoc, typeless instance.
> I don't want a custom type where I have to write boilerplate around for every single operation
That's why I use a language which lets me abstract away boilerplate:-)
> S-exprs by themselves are just data, they're the very same as JSON (not versioned, not extensible). YAML supports tags making it easy to plug your custom types directly into its reader.
I prefer (no surprise!) the Lisp reader, which does offer ways to plug your own types directly in with read-macros.
The dictionary will preserve the entirety of the data, that's already better than any custom type can ever dream to do. You're not going to want to upgrade your codebase every time a 3rd party adds fields to their documents.
Trying to make data "safe" by giving it strict types is only going to result in the destruction of data and no actual gains in safety.
> I prefer (no surprise!) the Lisp reader
The Lisp reader is not a data format nor is it available outside Lisp. That's no good on the wire unless you live in a sandbox and never talk to the real world.
You're going to receive data with fields/types your program do not understand (the endpoint could've pushed a new version without your knowledge) and there your ad-hoc types will become pain points.
Imagine if we tried to make HTTP "safe" by giving types to everything. The second a proxy sees a header it doesn't understand the request would fail. We'd have absolutely no internet. Distributed systems depend on the fact that you're able to handle unknown data and types.
If I need to work on the data I'll parse it into a data structure. Then I get something like ["Hello " [:b "World!"]] where I can leverage the existing functions operating on data in order to manipulate it.
With some formats like SVG you even need to parse the document into a DOM and work on it to get things right.
Lots of "features" that actually make things just harder to get right. Like XML entities can lead to a billion laughs [1] or sensitive files being included [2]. XML 1.0 files can contain null bytes, 1.1 files can't.
[1]: https://en.wikipedia.org/wiki/Billion_laughs [2]: https://www.owasp.org/index.php/XML_External_Entity_(XXE)_Pr...
For example, many clients (like mcabber) had lots of features broken (e.g. typing notifications) when interacting with GTalk because they couldn't parse XML namespaces properly.
Unfortunately, this disqualifies a lot of the less maintained and/or less well-written parsers. It also sometimes disqualifies "bindings" to things like libxml2, when the bindings are written by someone who doesn't understand namespaces and end up wrecking them even though the underlying library supports them fine.
XML namespaces are strange to me. I think they're the simplest example of a standard I consider somewhat simple that to a first approximation nobody understands. It really goes to show that there aren't really all that many people who understand XML qua XML. There's a lot of people who understand "there's angle brackets and attributes and sometimes these entities" and not much more, but there's substantially more meat to correctly understanding XML. Unfortunately (and I mean this fully, not sarcastically or with preaching), to use and implement XMPP requires a somewhat more full understanding of XML than most people have. Namespaces are the biggest sticking point, but XMPP is also a bit quirky in that it tries to have an XML stream, which is sensible enough at the protocol level but often translates strangely up to the API level. (The optimum library for XMPP permits you to use a SAX-like parser for the stream opening, then presents a DOM-like view of each individual element, and there aren't very many libraries that allow both sides of that.)
-- Data model --
1. Every tag and attribute is optionally namespaced by a URI. (Tags usually are, attributes usually are not.)
2. The meaning of a non-namespaced attribute is defined by the tag to which it's applied.
3. The meaning of a non-namespaced tag is defined by the application.
-- Serialization --
4. To save space in the textual representation of XML, namespaces are (necessarily) referenced by user-defined aliases.
5. To save even more space, a default namespace may be (and usually is) defined for use by tags (but not attributes).
6. Alias definitions live in opening tags, and are scoped to the tag itself and all its children. They are of the form xmlns:{alias}="{namespace}", or xmlns="{namespace}" for the default namespace.
7. Tag and attribute names are prefixed with an alias to indicate that they live in that namespace: {alias}:{tag/attribute name}
8. If a tag name is not prefixed with an alias, it lives in the default namespace (if any).
9. If an attribute name is not prefixed with an alias, it is not namespaced.
10. The aliases generally have no semantic meaning, but there are wacky exceptions (e.g. XML Schema).
EDIT, forgot: 11. The "xml" namespace is reserved and predefined as 'http://www.w3.org/XML/1998/namespace'.
Secondly, when you need namespaces, they're a clean, simple solution to the core problem. When you don't need them, don't use them. Fortunately, unlike some of the DTD stuff which is hard to ignore correctly, it's really easy to not use namespaces. If the complexity comes from the problem, trying to solve it with a too-simple solution is not a virtue. You'll just reinvent the same thing, only worse, and with less library support.
For proof... see the endless poorly-supported attempts to do anything even remotely similar for JSON. The good news is that JSON is very simple. The bad news is that when it's too simple, you typically end up with a mess of poorly-designed stuff on top.
Given that you find three semantic rules to be "too complicated", I suspect you'll have a difficult time doing better.
Because XML puts the "type" of an object syntactically first, it supports this use case.
[Edit: rephrase.]
But yes, it's an uphill battle if you don't have a namespace aware XML parser.
The thing that does _not_ work is not using namespaces to decide on the meaning of tags/attributes when processing XMPP :)
It took me 10 minutes to learn the YAML format and I find it incredibly more expressive, powerful and readable than XML.