Enabling the Future of GitHub's REST API with API Versioning
github.blog
github.blog
There’s zero practical user benefit to NOT specifying the version, so why not enforce it? You’re creating a footgun otherwise.
Furthermore as others have pointed out, this is a purely cosmetic change. Using /api/$UNIQUE_STR in the URL vs. X-GitHub-Version: $UNIQUE_STR in the header are functionally equivalent, so…
Why bother making this change then? What was suboptimal about the previous way? What benefit does this bring?
EDIT: Serves me right for not reading the docs more carefully. From the docs [1]:
Requests without the X-GitHub-Api-Version header will default to use the 2022-11-28 version.
See phphphphp’s reply below for context.
[1] https://docs.github.com/en/rest/overview/api-versions?apiVer...
I work with a lot of APIs that just have “/v1” or “/v2” as part of the endpoint so it’s easy to specify. Setting a header parameter to a specific date is a PITA and I certainly don’t want to look up the curl syntax whenever I do simple work.
I think if they made it required, it would annoy people. So they didn’t.
And it would annoy me far more if my app that talks to GitHub inexplicably breaks in a few years because of a release I wasn’t aware of.
These changes make no sense to me.
You can update each endpoint in non-backwards compatible ways without having to update all of them.
If you have it in the URL, it's likely libraries and custom scripts will just set the value in configuration instead of per-callsite.
You don't want to have to wait to fix up some API or make non-updates on all the rest just to release a new version.
edit: I agree that making it optional is a bad move
Here specifically I was suggesting that the /v2/ in the URL would make people using the API more likely to set the /v2/ in their app to use it across all calls rather than different per endpoint versions.
It looks like the calls are in the version rather than the version being in the calls.
It's a general problem I see with any REST API that tries to version that way. Version isn't a strict hierarchical "owns" relationship and embedding it early in the URL sort of violates REST principles in what the folder path of a URL is meant to imply. (Admittedly, that's a bit of a more strict interpretation than most people pragmatically follow when building REST APIs, but it is an interesting and useful strict interpretation so worth bringing up and examining.)
If you are versioning individual endpoints it might make sense to use URLs like: some/endpoint/v1 and some/endpoint/v2
That makes it more clear that "some/endpoint" "owns" that relationship and kind of/sort of what "exactly" is being versioned. If you must put a version number in the URL.
That said, REST has always been about content negotiation as a long held tenet, and that has always been about using the right Headers (since the earliest HTTP 1.x apps), and I've always felt like complaints that you can't send the right version Header reflects primarily on bad tools that don't understand HTTP (and REST) as well as they should. Sending Headers is an important part of HTTP. I don't think it should be harder to send the right Headers than to embed things that would be Headers into the URL. Though that's just my somewhat strong opinion on the subject.
But now I think that the versions are not part of the resource path, just as starting the path with /api doesn't mean the API is part of the resource. The versions are part of the api path. If I had a URL that was /api-v1 and /api-v2 it would be more obvious, perhaps, that the URLs are addressing different API implementations, and everything that came after was the RESTful bit. Having /api/v1 and api/v2 as the prefixes isn't quite as clear on that point, for sure.
The reason I like having the version in the path is it makes it a bigger deal for the implementer when they're making breaking changes. Having a per-endpoint, slightly hidden way of versioning endpoints might mean it's harder for them to intuit quite how many changes are flying at their API's consumers.
I think you are almost assuming the implementer doesn't abstract the URL routing in some way and has to care about about every possible URL route? In practice there is often very little difference on the implementer side between routing `/api/v{id}/some/route` and pulling in the version number as a route parameter and picking up a version number from a header. The implementer is still just seeing a version number parameter injected either from their routing framework or their header dictionary. In both cases the amount of code they see is roughly the same. In many cases all "versions" of the code might only be a single "Controller". Outside of a few languages such as PHP there's few languages that have a direct 1:1 between URL paths and "Controllers" or "Source Files" or anything resembling that. At some point there's always some intuition that side-by-side versions imply side-by-side code and the "tech debt" sense of that maintenance version, but that intuition comes from numbers of classes and size of classes and indirect measurements like PR cycle time and Sprint planning estimates, and is never directly connected to URL routing versus Header negotiation. (Again, in modern languages with modern routing frameworks. Obviously if you have to drop a PHP or ColdFusion file for every API endpoint you'll feel URLs a lot more painfully than that.)
Really? Don't you do this all the time for everything beyond a simple GET? From my experience it's far preferable than URL versioning because you can version on an individual endpoint. Right now I'm consuming Versions 0.9, 1.0 and 2.0 from a vendor; it would be way nicer to set individual headers than somehow mananging multiple endpoints.
That being said, I don’t call multiple versions of an api so maybe it’s nice to just be able to call multiples.
I imagine that it’s pretty easy to manage with the url version as you’re still just storing one endpoint and appending the version depending on what you would like. So programmatically you just change a config variable with the version you want.
But it seems like the version on changes with breaks so if you’re calling different versions you have to handle the input and output differently anyway.
I don't understand the difference. What makes it nicer?
I did a quick search and couldn’t see any docs on this use case.
And it behaves the same way as `curl --header "" --header ""` in that it's additive:
$ curl -vK <(printf 'header = "alpha: 1"\nheader = "beta: 2"\n') httpbin.org/getThe benefit of this new versioning approach is that they can maintain backwards compatibility for existing implementations while introducing breaking changes for new implementations to benefit from.
In an ideal world, sure, they should have implemented this from the start, but they didn’t, so going forward they have to accommodate both existing implementations (by not breaking anything) and new implementations (by providing features that would break existing implementations).
They’re basically just implementing stripe’s well-understood and battle tested approach for versioning, while maintaining backwards compatibility. I am struggling to see any problems with the approach, in fact, it seems like the only right approach (?)
That certainly addresses the footgun I mentioned.
EDIT: From the docs: Requests without the X-GitHub-Api-Version header will default to use the 2022-11-28 version.
At some point the default 28th Nov 2022 version is unsupported... So what is served must change... Either it's the latest at the time or the behaviour of receiving a response without providing the header is removed.
I'll take the bait, like C++?
> comfortable making incompatible updates
Incompatible how? You specify the version, your semantics are of that version.
> When a new REST API version is released, we’re committed to supporting the previous version for at least two years (24 months).
The blog post suggests that we might find out some of the API changes that warranted this change as early as next month, but it would have been great for this blog post to include some examples of changes that were coming.
I don't think it's a strong point.
As someone who has written some fairly intensive github integrations, I think the biggest thing github could do DX wise is unify the systems behind the two APIs and make them actually have parity. I'd always need both because they are a Venn Diagram, feature-wise.
Though to be fair, I'm guessing it is a very difficult problem that every engineer on both teams dwells on regularly.
I frequently perform basic API exploration in my web browser - hitting https://api.github.com/repos/simonw/datasette/commits for example - and there isn't a convenient way to add headers to those requests.
I quite often do this kind of exploration on my iPhone in Mobile Safari (no need to get my laptop out just to quickly check what shape an API is), which is even less convenient for sending custom headers.
The proposed scheme is fairly similar to Stripe's, and although Stripe mainly uses the `Stripe-Version` header for versioning, its API has also always allowed you to alternatively send a `?_stripe_version=<version>` query parameter which it'll respect.
Stripe's API implementation definitely had its share of legacy cruft that added up and made things more difficult to maintain, but out of all of it, I don't remember having to support that alternative `_stripe_version` parameter ever really being a maintenance headache. The leading underscore made sure that even as the API expanded, it never accidentally collided with parameters on any other endpoint.
That being said, it’s not the end of the world. It’s just eventually these straws will add up to something really bad.
I definitely see the benefit of being able to easily hit the URL from a browser - although I think it's probably only relevant for unauthenticated GET requests.
For the last 12 years I have put the api version in the url, it is unambiguous and has never caused a problem. In-band is best, esp for a public service that is consumed in an ad-hoc manner (wget, curl, any http library you can find). Hell, even include the docs in the response.
Hit https://old.reddit.com/domain/simonwillison.net/.json in Firefox and you can get a long way.
An alternative answer to this question is: well, if it's basic then you don't care about which version they respond with :-D
Semantically it is equivalent though.
Sure if I’m writing an integration that’s no big deal, but it makes it more of a PITA to do things from the CLI with curl.
that would be my guess
In essence, they transform requests from the latest API version back to the (older) versions by writing an adapter for each response for every breaking change. Not sure what they do for new endpoints, or when a future response no longer includes a field that used to be included before.
Some nice things though:
- timeline they suggest seems to be longer than GitHub (implied by the “a power company should not change its voltage every 2 years” statement)
- instead of a header, it’s linked to the account of the user and the time of the first API call. That might be a good default for GH rather than the oldest available version at that time as suggested elsewhere in the comments. (drives adoption, least surprise, longest expected validity of the interface)
Edit: formatting
I personally find this a much saner way to declare api versions (version in header). Specially if the platform supports Webhooks which have urls in the payload: if your API declares version in the path, which version do you use in the webhooks urls?
That seems really surprising.
It seems not strong enough for Microsoft. But that’s why I was surprised as being a not very useful hassle would seem to be a big reason not to do it that way.
Cynically, I fear this method allows lots more versions and lots more breaking changes so get ready for stuff that could have been backwards compatible now just be a breaking change with a new version.
But the URL version inly wins that argument if you ignore all the issues.
Github clearly wants to iterate their REST API faster and feels constrained, the URL method would mean either the entire API versions over quickly with almost but not quite copy pastes, or endpoints get versioned individually and you’re simultaneously hitting /3/foo and /42/bar with little rhyme or reason.
Things get confusing when you retrieve an entity via /6/ then update it via /4/.
Most api consumers don’t care if fields are added or new methods added they are not using, thus they don’t care much about the version unless upgrading. Removing things or changing behaviour should be very rare and have a very good justification.
I am of the opinion URLs are a promise the document will be available until the end of time. I don’t want to break the internet with 404s.
Since the only promise here is for the API to exist for 2 years when depreciated, then I think using a header is appropriate.
Query param doesn’t make sense either as it is often used as input data into the document at the URL. The document schema depends on the API version, so query params become a chicken/egg scenario
Short version: I do initial research against APIs using web browsers a LOT, including on my phone, and headers are really inconvenient for that.
In case you can't see DEAD comments, there's one here that talks about this too (and isn't offensive about it, so I imagine the account is DEAD for some other reason): https://news.ycombinator.com/item?id=33781215
Let’s say I start coding against the API today, can I just set the header to “today” and github will dereference that to the nearest preceding version, or do I have to look up and set the header for each endpoint?
IME it’s common to have a “broker” which handles all calls to the API in order to do common ore and post processing, having to look up and specify versions for each call sites would be frustrating, as well as the risk of different calls using different versions.
Imagine you release package on npm (or whatever) versioning it with v1 and constantly adding new features and still keeping the v1 version tag. That's what you're doing to your API.
If the version is part of the URL, you have benefits such as -
1. It's easy to inspect the URL. For headers it's more involved.
2. Middleware (Proxies) can drop headers that they don't recognize.
3. Logging frameworks typically log URLs but don't log headers OOTB.
As someone who has built a lot of APIs for a long time, I think the decision to use headers to communicate versioning information is not optimal.
> In the end, I decided the fairest, most balanced way was to piss everyone off equally. Of course I’m talking about API versioning and not since the great “tabs versus spaces” debate have I seen so many strong beliefs in entirely different camps.
https://www.troyhunt.com/your-api-versioning-is-wrong-which-...
Edit: it's been fixed now, good Github doc fix response time!
That said, the overall concept is deprecated: https://www.rfc-editor.org/rfc/rfc6648
So if your header is something you think should/will be standard but it's still in beta/testing, give it the real name.
Please also return error if the wrong key was used, instead of silent fallback of assuming no version.
Blogs/email shouldn't be the only channel. Announce the retirement date with an HTTP header right in the response. In this way, developers can create automated alerts about soon depreciation.
In other words, depreciation announcements should be machine-friendly, not just user-friendly.
That said, as long as there's multiple ways to get notified, more is probably better.
As @xavdid noted, the reality is that most people won't see these headers, but it's a nice power user feature.
It is a little surprising that the version defaults to the 2022-11-28 rather than defaulting to latest if omitted, but I guess that's a way to err on the side of not breaking old users when they upgrade an SDK.
Just so you know, we don’t intend to deprecate `2022-11-28` in two years-ish - we plan to keep versions around for much longer than that. But we offer a commitment of at least 2 years for peace of mind.
You might consider behavior more like Stripe's here (disclaimer, former member of the stripe API platform team).
As you're likely aware, Stripe defaults to the latest version for each user, and then stores that as the user's default API Version until they upgrade it at the account level.
This works pretty well, and overall I'd recommend it. It may be better to give each API Key its own default API version, since different applications may expect different versions, and API keys can be a simple way to differentiate applications.
Some of Stripe's client libraries also enforce sending a specific API version header to ensure the types are correct; my guess is that in practice, this is how most requests have their version dictated now, making the account-level setting a little anachronistic.
Everything gets trickiest when it comes to webhook bodies, since the way the payload is rendered may be different across versions, so you can have one app create an object on version A and then a different app consume the webhook under version B and see the created object with a different shape. I'm not familiar enough with GitHub's API to know how best to handle this sort of thing.
Anyway, good luck and happy to chat more.
That would not work if I (leipert) build a client against version 2022-11-28 and you (rattray) create an API token to use with that client on a later date though, right?
[*] never meaning a long time
I found media types convenient in that they can be explicitly specified in an OpenAPI specification.
Honestly, we found that media types within the GitHub API had just got pretty confusing as we use them for so many things - "feature previews", response formats, versions (https://docs.github.com/en/rest/overview/media-types?apiVers...). We didn't want to overload them anymore.
Just so you know, we published OpenAPI specifications for all of our API versions on GitHub: https://github.com/github/rest-api-description/tree/main/des...
I know it's not that REST .. it is what it is, probably annoyingly strict req/res, but it's very solid API. I'm glad they are trying to improve .. it.
We’re just aware that GraphQL is pretty different, so if we want versioning there, we’ll need to tackle it differently. (On top of that, GraphQL has some nice primitives that come for free, like deprecated fields.)
i missed this, is there a canonical source to learn more?
Diff format example: {added: [{"path-to-key": {..}}], removed: [], changed: [{"a/b/c": {prev: {type: "string"}, current: {type: "number"}}], moved: [{old: stuff1, new: stuff2}]}
Supporting many API versions in parallel definitely has an overhead from an engineering perspective - but, as the blog post says, we thought it was crucial that we unlock the ability for us to evolve our API whilst still providing consistency to existing integrators. Some form of versioning was the only way we could see to do that.
We've built some pretty neat tooling and abstractions internally to make it easy for teams across GitHub to build and test multiple API versions. We hope to publish a blog post about that in the near future.
We have some small changes queued up, but nothing huge or earth-shattering.
Our focus right now is on improving developer experience and consistency where we have some weirdness in our current API design that trips integrators up. Without versioning, we couldn't address those papercuts.
We plan to release the first breaking version in the next few months.
I would be remiss if I didn’t take this opportunity to ask: the current API for retrieving commit metadata relies on git’s information about the committer, but you have much better information available to you. You know who actually pushed the commit.
In an enterprise setting, that’s a big deal, because it’s easy for users to have misconfigurations. We have no shortage of commits in our repositories from ec2-user or root.
I don’t think it makes sense to replace the information git supplies, but could you please augment it?
I don’t think companies have to support forever but I don’t like the idea of having to rework my integrations every two years.
Also specifying a date is very clunky. I’d rather they use semver (like a non-barbarian) and stick with breaking changes going into a major release.
Or how many releases are between 2022-11-28 and 2042-11-28? I’m sure there’s a docs page spelling it out with details, but being able to compare two versions and tell if your code will work is pretty handy.
How much the major version changed (v2 -> v3 vs v3 -> v9) doesn't actually tell you how many changes there were. In that way it's almost identical to using major-version-only semver, vs dates.
Consider v3 has been around for a decade, they seems likely to last longer than 2 years.
If they intend to maintain them for 10 years then don’t set a specific timeframe and just stick to whatever their current nebulous language is.