Using JSON Schema to Document, Test, and Debug APIs
blog.heroku.com
blog.heroku.com
[0] https://swagger.io/docs/specification/about/
We've got some JSON Schema APIs in our services already and are using libs like the Python `json-schema` package to do validation with those. I also briefly experimented with Quicktype [2] to generate some TS types as a proof-of-concept.
We're getting ready to do a pretty big rework of a lot of our services, and I'm interested in any info folks can provide on pros and cons of using OpenAPI vs JSON Schema for API definitions, and tooling around request validation and TS interface / client generation.
[0] https://philsturgeon.uk/api/2018/04/13/openapi-and-json-sche...
[1] https://github.com/OAI/OpenAPI-Specification/issues/1532
[1] They don't want to call it a fork, but when they invent totally new extensions to support things that already is solved by parts of the standard that they don't want to support, then it is a fork.
OpenAPI 3 was developed while JSON Schema progress was stalled due to the prior editors leaving and a new group of us (eventually) picking it up.
OpenAPI 3.1 will most likely include a keyword to allow using standard JSON Schema as an alternative to their customized syntax, and hopefully we can achieve full integration on OpenAPI 4. There are also some other ideas being explored for improved compatibility in 3.x.
Standards work is hard, but the relationship between the OpenAPI TSC and the JSON Schema editors is quite healthy and we are making good progress.
Not too long ago, I tried to build a project around OpenAPI, trying to generate model structs in Go from OpenAPI definitions. The support just wasn't there. There were code generation tools for v2 and emerging library support for v3, but nothing that covered both.
I went back and used plain JSON Schema for my models, and used GraphQL instead of OpenAPI for the API, and that turned out great.
Also, we and the OpenAPI Technical Steering Committee are actively working together to re-converge the specifications. OpenAPI 3 was developed while JSON Schema progress was stalled due to the prior editors leaving and a new group of us (eventually) picking it up.
OpenAPI 3.1 will most likely include a keyword to allow using standard JSON Schema as an alternative to their customized syntax, and hopefully we can achieve full integration on OpenAPI 4. There are also some other ideas being explored for improved compatibility in 3.x.
Standards work is hard, but the relationship between the OpenAPI TSC and the JSON Schema editors is quite healthy and we are making good progress.
It seems like Swagger v2 was reasonably well-liked so the idea was to make OpenAPI the best possible way to describe any kind of REST API in existence. Since the spec got more or less merged with JSON Schema it has become extremely unwieldy and full of edge cases.
At my company we're trying to develop tooling around OpenAPI but a lot of things (especially related to the 'oneOf'/'anyOf'/'allOf' features) are extremely ambiguous.
Not to mention that at least the Java versions of the parsing libraries are full of inconsistencies as well. For instance, there's a OpenAPI parser which also has a 'compatibility' mode so it can read Swagger V2 specs as well. Unfortunately the data you get from reading a Swagger V2 spec via that compatibility layer is different from when you'd first convert the V2 spec to OpenAPI v3 through an external tool.
The project is certainly ambitious and I completely agree that this is a hard problem to solve, but honestly I would say that if you want universal adoption of a toolkit for writing API's then the API specification language itself should not be so difficult or ambiguous in its implementation.
It took a lot of work to get there, and there's still lots of room for improvement. My ultimate goal is to automatically generate the API client test suite based on the requests/responses. I also want to add some custom logic and workflow rules to the OpenAPI specification, instead of needing to write custom wrapper code in 10+ languages.
I've had to write the same basic code so many times: Make a POST request, then make a GET request once per second until the "status" changes from "pending" to "done", and then finally return the completed result. (I know it would probably be better to set up a websocket connection, but this works fine and it was easier to implement.) It would be so awesome if I could define this workflow inside my API specification, and then the auto-generated client code could include this polling logic without any effort on my part. (Apart from needing to write and support the higher-level generator code for each language, which would actually be a lot more work.)
The other really annoying thing is needing to figure out package managers and release steps for every single programming language. I've only figured this out for 6 languages so far (C#, Java, JS, PHP, Python, Ruby), but I want to support far more languages, and it's just exhausting to go through this process each time. It would be so nice if there was an open source project that wrapped all of the different package managers and provided a framework on top of openapi-generator. And if there was a CLI tool (or web UI) that walked you through the process of signing up for accounts and setting up API keys, and then keeping it all in one place. I would honestly be tempted to rewrite openapi-generator from scratch in a better language / template engine, because the Java code is so hard to read and extend. I'm mainly a Ruby developer, but I don't think Ruby would be a great choice for CLI / generator tools, because it gets pretty slow. So maybe Python or Go.
I feel like this would a really interesting project to work on, and potentially a startup idea. Would anyone be interested in using it? I should probably try to validate this idea. I've set up a Google Form where you can just click a checkbox to register your interest anonymously, or you can also submit your email if you want to get updates and try a beta version: https://forms.gle/7qWzWpC9QrTjUgnU7
The workflow which I use for one app is:
* Define a JSON Schema for all the app's models
* Generate static Go types from these definitions
* Generate corresponding GraphQL types and inputs in the Go app for serving the API
* Generate TypeScript types in JavaScript front ends against the same schema
* Same model structs in Go are used to shuffle data in and out of data stores
Since GraphQL is more limited than JSON Schema (for example, very limited union ("oneof") support), there are some features I'm not able to fully make use of, but the common denominator covers most use cases.
I love having end-to-end static typing, validation and consistency.
Can you link to the libraries you used??
After the point they added all logics to the spec, everything becomes a mess.
Example 1:
{ "oneOf": [ {"type": "number"}, {"type": "number"} ] }
1 or 2 or 3 is an invalid input based on above schema. It doesn't make sense to me at first glance. (hint: it's XOR)
---------
Example 2:
{ "oneOf": [ {"minimum": 0, "maximum": 10}, {"minimum": 5, "maximum": 20} ] }
These ranges will pass: [0,4] and [11,20] But you would be surprise it doesn't reject string, bool, null. Basically it doesn't reject anything except [5,10] range.
---------
Example 3:
{ "type": ["object", "array", "null"], "not": {} }
Easy but confusing. This rejects every thing because comma means AND, and `"not": {}` means false.
---------
Example 4:
{ "allOf": [ { "type": "object", "properties": { "a": {"type": "string"}, "b": {"type": "integer"} } }], "additionalProperties": false }
A well-known problem. This rejects (all - {})
This is because the "allOf" is object AND outside it is an empty object. "properties" at level 0 is empty if not specified. So this schema only accept {}; an object AND is empty.
They add a new keyword called "unevaluatedProperties" to solve this and I won't explain it to you!
-----
If you just use basic stuff like draft-04, you will be fine.
But I will NEVER touch this spec again!
-----
EDIT: My app was sort of static analysis on schema, and this spec doesn't suppose to help doing thing like that.
> "oneOf": [ {"type": "number"}, {"type": "number"} ]
> { "oneOf": [ {"minimum": 0, "maximum": 10}, {"minimum": 5, "maximum": 20} ] }
"oneOf" requires exactly one match (hence the name), typically you want to use "anyOf" or "allOf". And there's no benefit in putting two identical values inside them.
> { "type": ["object", "array", "null"], "not": {} }
> { "allOf": [ { "type": "object", "properties": { "a": {"type": "string"}, "b": {"type": "integer"} } }], "additionalProperties": false }
JSON Schema is merely a list of assertions. Some test the type of the value (that's the "type" keyword), others test the value ranges within a single type. This way, you can allow values to be one of multiple types, e.g.: {type:["string","object"], minLength:1} means "Value must be a string or object; and if it's a string, it must have at least one character."
Some of the assertions are spread across multiple keywords, "additionalProperties" depends on "properties", for example. So, {additionalProperties:false} means: if value is an object, then only an empty object is permitted.
Some implementations can tell you if you're trying to do nonsensical things (like test the maximum length of a value that's only allowed to be a boolean), but that's up to the implementation to test for.
The fact that the spec is too flexible that allow users to write all nonsensical things, and so sensible things that might have hole on assertions.
I understand all your explanation and thanks for creating a new account to do this, appreciated!
Also this: https://www.genivia.com/sjot.html#SJOT_versus_JSON_schema (JSON Schema problems)
GOOD LUCK!
To me it seems the XML era had better engineers making the standards.
There is still some performance cost for checking a response against a schema though, so a nice compromise is just to check a small % of outgoing responses — the vast majority of requests stay as fast as possible, but given a reasonable traffic load, it's still enough to eventually reveal any places the schema doesn't match reality. (This is an approach we use at Stripe for checking our OpenAPI specification.)
Currently we are checking 100% of the responses. I even wrote this gem which uses native code to perform the schema validation to minimize the overhead in our endpoints: https://github.com/foxtacles/rj_schema Validating our largest and most complex responses is taking <10ms, on average no more than 2-5ms which is quite affordable.
I tried using JSON Schema for some internal (non-HTTP) APIs and found bugs/inconsistencies in tooling that made me abandon it.
pro: It validates API JSON responses based on an open spec.
cons: It's absolutely terrible to work with. XML all over again. Documentation is terrible. Finding examples is, well you get it...
Why are we trying to shoehorn a schema design language into a format that isn't good for human authoring? It's annoying to write JSON, yet the language masquerades as being human-readable.
Frontend engineers should check out protobuf. It's an amazing data definition language that generates bindings in every language under the sun and has a compact binary serialization that is much more efficient than JSON both to encode/decode and transmit over the wire.
JSON should die the same death XML did. It's so bad.
For example oneof customer (individual / organisation) with different distinct fields. In protobuf, all these fields need to be defined.
Even Hyper-Schema (to the extent that it's implemented at all yet) is a resource-by-resource system, not an API-scope system.
What's wrong with simple https://github.com/omniti-labs/jsend
``` { status : "success", data : { "post" : { "id" : 1, "title" : "A blog post", "body" : "Some useful content" } } } ```
What part of the article is verbose? What's wrong with verbose?
It's not perfect but a great time saver for me and users.