Robust APIs Are Weird
aviskase.com
aviskase.com
[1]I guess if you really stretch things, you can say 1 is false and 0 is true, and I have actually seen that before, but I can't think of any other ambiguities a single bit could have.
Adding to that line of thinking: obviously we need the textual representation for reading - so we need something to translate it. That situation is kind of dual to programming languages and compilers - we use compilers to make the textual representation machine readable, here we need a decompiler to make the machine readable data textual.
But if you start from a programming language, it surely do have something to describe data structures. That would already be unambiguous. That is why JSON is useful.
The text-based language YAML allows “yes” and “no” for Boolean in addition to “true” and “false”. So if multiple bright people think about this topic they’ll evidently come up with different approaches. Not so with a single bit.
Oh no... be very strict in what you accept. Once out there, you can't make it more strict. APIs should be as strict and defined as possible. Unless the API is not an API, but user input, then, yes, be liberal in what you accept
In a modern context this would correspond to ignoring unknown keys in json, or ignoring unknown fields in protobuf.
Sadly many interpret this as accepting an ill defined superset of the specification, which then turns into a complex specification as implementations rely on those extensions being handled in a specific way.
If you want to use json, validate, and accept both 10 and "10" you can use json schema and do:
``` { "type": ["string", "number"]} ```
Depending on the library and language you could extend the rules to accept a string that can be parsed and throw if not.
You can use it to validate your JSON inputs, much like you would with XSD (I've used both).
There is OpenAPI, née Swagger [2], which allows you to describe your API in human- and machine-readable form and generate code for it [3].
But validation is but one step. A robust API needs rate-limiting, injection protection (does your language has tainted strings?), easy enough authz/authn story, etc, etc, all different and non-trivial things.
[1]: http://json-schema.org/ [2]: https://en.wikipedia.org/wiki/OpenAPI_Specification [3]: https://openapi-generator.tech/
XML schema is a big reason that I would use XML. Otherwise, it’s a fairly painful and prolix standard.
XML is also a good way to deliver longer streams of data (for me), because so many parsers will allow “realtime” parsing of “broken” XML data, while JSON tends to require the entire document to be delivered (and correct) before parsing (I’m sure that will be changing, if it hasn’t already).
I’ve written ONVIF software, which is SOAP/WSDL-based.
I’ve written XSLT, which could be used as a form of medieval torture.
I’ve written servers, with REST, and REST-like APIs. I generally prefer REST-like, where the server response is JSON and/or XML (usually either one), but the GET/POST/PATCH stimulus is transaction/URI argument-based. That’s the way most folks seem to like doing it, as well.
Pure REST requires sending XML or JSON; putting a huge onus on the server, and making it difficult to be flexible (I like to support both XML and JSON as “first-class citizens”).
> Because the simplest way to validate a JSON document is first to consume it with some common library that guesses types almost like we do: “does it have quotes around? string!” And only then, with already cast value, to compare its type with whatever is defined in a schema.
This is true. JSON schema isn’t particularly useful to me. I feel like it works against the reason for using JSON. I like JSON, precisely because it is so lightweight.
XML Schema is incredibly mature and robust (and an enormous pain in the butt to write). Validation is built into almost every parser; while many folks are unaware that JSON Schema even exists.
What I tend to do, is serve both XML and JSON from a server, with XML accompanied by Schema (sometimes, dynamically generated), but the JSON derived from the XML, so it benefits from the XML validation.
I’m wondering if standards like OpenGL will replace both XML and JSON. I’ve never used it, because of the requirement to use a third-party library, but it is a compelling tool.
The way I do it is to define classes in C# and annotate with attributes (either DataContract and/or System.Xml.Serialization) and then export the schema as a set of XSD documents with the accompanying catalog file.
Sorry. My bad. PBC (Posting Before Coffee).
GraphQL
"JSON lines" extends JSON by a) not allowing newlines in values and b) allowing newlines to be used as a top-level separator for elements of a list, instead of the JSON-compliant comma-separated list enclosed in square brackets
This allows for per-line parsing, which can significantly reduce memory usage and allow for partial parsing in some scenarios. I've seen it used where the data is fundamentally a large collection at the top level where producing proper JSON was using too much memory
There should be also load balancing/HA, rate limiting, authn/authz, injection, all sorts of protections against OWASP top10 type attacks.
Linking to the right time: https://youtu.be/djKPtyXhaNE?t=1128
TL;DR:
- GraphQL is a descriptive type system not prescriptive. This means that "String" maps to JSON String maps to Javascript utf8 string, but what type that maps to for C++ is not quite clear. Because there are many different ways to represent a string. So which one do we use? As opposed to protobuf client codegen which is prescriptive and unambiguous in what you get.
- Encoding custom scalars can cause some confusion because it's not clear what the spec of the custom scalar is. Although, there's been some recent work to make that a little better.[1]
- Impedance mismatch with nullability and union types because support is not uniform across languages
First of all, I agree with all the comments mentioning that being robust implies much more than validation. This article was a take on one example I found curious.
And second, as you may notice, that article was for testers. That's why a certain oversimplification.
Personally, I am on the side of "if the contract is defined, follow it." All the fiddling with supplied data seems fragile and prone to hidden behaviour. But there is a valid argument for cases when your data suppliers have a history of ever changing and/or buggy output. If you pay them money, you can request them to fix it. If they pay you money, perhaps you would consider being more flexible.
The downside is that you need a special tool to accesss it, a plain browser or curl won't work.
You might be thinking grpc-gateway. [1]