> But saying "it's just the same as JSON" misses how many dangerous parts and footguns it has.
I didn't say it's just the same as JSON. I said:
"And as long as you’re not allowing users to upload their own XML, then you get to control the schema so there isn’t any risks in using XML."
Every example you've given requires untrusted 3rd parties to craft the XML. But that wasn't what I was advocating here. I was talking specifically about the API returning XML.
> and give it to few different XML programmers, each one of them will come up with a different schema.
But again, the API is controlling the schema so this isn't an issue for the use case I discussed.
> Compare to JSON, where this can have only one canonical encoding, and the worst you might have to deal with would be some uppercased letters.
You've clearly not worked with enough JSON if that's all you think the issue with JSON is. I've written JSON parsers and used a fair few open source ones too. And there's a lot of places things can go wrong:
1. You have number serialization bugs between different JSON parsers.
2. No standard for dates. Causing everyone to do things slightly differently
3. Inconsistencies with top level arrays, some parsers require top level arrays to be `{[ ... ]}` whiles others are happy just with `[ ... ]`
4. Parsers don't all agree on how to represent non-alpha / numeric ASCII characters. And we're not just talking about unicode, Even some ASCII characters like `>` can be handled differently by different JSON libraries
5. Lots of different JSON supersets (because JSON itself doesn't support half the stuff that people need from it), like jsonlines, concatenated json, newline delimited json, JSON with date fields (as seen in popular JSON libraries in .NET), JSON schema, etc.
6. Even your key name example has numerous other inconsistencies you haven't touched on. Like UPPER, lower, dot.notation, hyphenated-keys, underscored_keys, UpperCamelCased, lowerCamelCased...and so many variations in between.
Ignoring JSON supersets, then I agree that JSON has fewer places for exploits in user generated documents. But the specification is also only 5 (FIVE!!!) pages long and thus it allows for a lot of undefined behaviour. And that's a problem for somethings who's entire purpose is a database.
This is why XML is so complex -- precisely because it's intended to solve these problems. But it was also intended to be served from trusted identities. Which is where the vulnerabilities lie.
> JSON is a simplified version of XML, and JSON has copied good XML features (it's not getting namespaces or external DTDs, and good riddance!).
I wouldn't be so sure about that: https://json-schema.org/specification
---
To go back to my earlier point: literally no-one is going to argue that XML doesn't have it's warts. But what you need to understand is that in the specific example that started this conversation, the API provider is the one defining the schema and crafting the XML. So literally none of your examples apply what-so-ever. In fact, this falls squarely under the correct usage of XML.
Context matters. User supplied XML is bad but that's not what is being proposed here. And that's why you're being called out of stating what you believed to be pretty obvious advice.