The RESTful CookBook
restcookbook.com
restcookbook.com
"Even though it's tempting to create your own pagination scheme by "building"
URLs, you should only use the information given by the API. This means, you
cannot assume that the 3rd page can be found at /collection/3. If you do,
your client will break as soon as the API changes its pagination behaviour."
Is that a legitimate caveat?Are there clients that don't break when you change the API the client works with? Pretty smart clients.
A lot of popular web APIs, e.g. Twitter, do not embrace HATEOAS, which is why API changes break clients. Which isn't really a big deal if you control all the clients.
But I guess it seems like neither of those are domains where anyone I know is really working; admittedly entirely anecdotal. API seem to change or get replaced by "the next best thing" on the order of months or years. And doing HATEOAS seems like making a c++ program const-correct.
Always starts out well but pretty soon isn't viewed as worthwhile by powers that be. Do you have an example of a really good HATEOAS public API that is widely consumed?
I would argue that cons-correctness is probably one of the most important considerations of any public C++ API, so not really the best analogy ;-). Try to refactor some actively used C++ framework API that never bothered about const correctness some day, weeks of fun and bickering with your API clients guaranteed.
The specific ones are first, last, next, and prev/previous. As long the name doesn't change, the URI behind the scenes should be able to change however you need.
We do this in the Twilio API though we don't use the suggested Link Relations to describe the URIs: https://www.twilio.com/docs/api/rest/call#list-get-example-1
Yes, you can write clients for GitHub's API that works that way, for example.
Here, I whipped you up one: https://gist.github.com/steveklabnik/5282209/e48b100d95ff77f...
This client will return you information about the first gist on the second page of the public list. It does this by fetching the list of gists from the root, then using the pagination to walk forward. GitHub can change all of this, and it'll still work.
$ ruby fetch_gists.rb
DATA:
ID: 5282131
URL: https://api.github.com/gists/5282131
USER: sparksp
Bonus: here's a slightly refactored version. The logic is more clear when you move things into helpers. https://gist.github.com/steveklabnik/5282209/52ed62366586838...I would suggest a better goal to be "get the 11th public gist". Although that still isn't perfect since it doesn't mention how the collection would be sorted (I presume by date).
1) Yes, you should only use info supplied by the API. This should mean that the API is helping to ensure the client never sends the user down a UX dead-end, and also that the client does not break when the API changes.
But...
2) I don't like the thought of collections having a resource identifier for a page. As items are deleted from the collection, the resource identifiers now represent a modified resource and break caching (delete item 3 of a 5 item per page collection, and all resources in the collection - logical pages - have been modified implicitly).
I much prefer using query string for this, as it is a query on a collection.
But... the API should still be the thing that generates/builds these URLs, the client should never do this work. The client must use the provided first|prev|self|next|last URLs, otherwise the client may break in the future.
An example of how I prefer pagination:
"collectionName": {
"total": 861,
"limit": 10,
"offset": 100,
"pages": 87,
"links": [
{"rel": "first", "href": "/api/things?limit=10"},
{"rel": "prev", "href": "/api/things?limit=10&offset=90"},
{"rel": "self", "href": "/api/things?limit=10&offset=100"},
{"rel": "next", "href": "/api/things?limit=10&offset=110"},
{"rel": "last", "href": "/api/things?limit=10&offset=860"},
],
"items": [
...
]
}
Where 'limit' and 'offset' are provided by the query string (but default to 25 and 0 respectively) and the API takes care of generating valid links (like not including 'prev' if you are on the first page').This way the client does what it should, just use the links provided.
And whilst I'm here... one of the things I dislike about link relations is how they presently fail to mention the method you should use (let alone content-types acceptable by that end-point).
In the example above every link is a GET, in the example on this page http://restcookbook.com/Basics/hateoas/ they're probably POST. In the link relation assignments http://www.iana.org/assignments/link-relations/link-relation... 'edit' is probably a PUT. There's an 'edit', but not a 'delete', so it's not as if just sending more than one 'rel' might describe the method.
Today the audience of an API is a developer, so it's fine to just point them at the documentation. But tomorrow it may well be a computer.
Mainly, I've found that for dynamic data sets, the links could possibly change between the first page and the 87th, as well as the number of pages, voiding the practicality of the data set.
Give them the number of records per page, the offset, the number of records, and let the clients determine the rest (ideally you'd have documentation showing how to paginate, whether it be ?page=1 or ?offset=20, or whatever).
That looks very complete though.
What is a client supposed to do when a new "nextToLast" appears? I've yet to hear of any API consumer that is driven like a browser.
What's the actual point, other than some purity to some abstract concept? Just make it part of the API to take /api/things?page=x and do the right thing.
What's "tomorrow it may well be a computer" supposed to mean? We'll have AI? Or there will be some sort of new WSDL "... for REST" invented?
The actual point is decoupling. It's generally considered in software architecture that design that's decoupled is a little harder to build upfront, but survives change over time.
As a 'real life' example, because Twitter exposed internal details (a tweet ID), clients sorted based on that number, because they assumed it'd be monotonically increasing. Once sequential IDs became a problem with Twitter at scale, they had an issue: if they switched to GUIDs, clients would break, because they were sorting based on the ID value. They had to invent an entirely new algorithm (Snowflake) to adapt to this change. This wouldn't have happened if they hadn't leaked internal details. It's basic encapsulation.
I can see how returning hypermedia adds yet another layer of abstraction (and potentially plenty more round trips!). I'm just unsure how it helps. I don't understand how actual client code (besides a browser) can deal with arbitrary hypermedia. I'm cautious when I don't understand why people are hyped up about something, but I've yet to see any "real life" examples that demonstrate real benefits of this approach.
There are always agreements between the API designers and client designers, because the code isn't intelligent, so it doesn't know what 'parent' means. If the name of that changes, then the client would have to change. What can be different is the url structure to get the 'parent'. Maybe there's a new token that has to be there, or the naming of something has been changed, or whatever. Those changes can be made without breaking the clients. Putting this in the responses means two things:
1) Small changes don't break everything (pagination is a great one, its a /pages/1 one day then ?page=1, then ?offset=25, then ?pageSize is optional, then it's required, etc.).
2) Fewer assumptions baked into the client. Getting the tweets for a user you've loaded would be nicer as load(user.tweets) than load("/tweets/"+user.id)
watch these two very good talks on this topic:
Explicit link relations aren't the only way to do hypermedia and the best practices are still moving/being discovered.
WADL is a machine-readable XML description of HTTP-based web applications (typically REST web services).
http://en.wikipedia.org/wiki/Web_Application_Description_Lan...Link relations can absolutely say this. See ATOMpub as one example, it's `edit` relation specifies that you should use PUT.
But, my point was more that it's a bit opaque. That there is a defined list of link relations ( http://www.iana.org/assignments/link-relations/link-relation... ) , but that list does not tell you the method to use.
edit and edit-media are covered by ATOMpub, but rel is not an attribute with a fixed range of values. Which means that the information a developer needs isn't there with the link, instead one has to go off and for each link relation try and find the original first citation, hopefully in a spec, which defines for that particular link relation the method to use. Then the developer has to hope that the API implementor also knew about that spec and did the right thing.
My criticism is that link relations fails to describe the valid verbs for a link. And there is no appropriate place to put it. We're being told about a link, but now how to interact with it.
We're given half of the story: this link relates to this item in this way (rel attribute), and you can expect it to hold this content type (type attribute) and it's over here (href attribute).
But the missing part of that story is "and to make a call to this particular end point you need to use one of these methods: HEAD, or GET".
We could use OPTIONS of course, that's the point of it. But pragmatism kicks in and when I've in the past gone done a purity route and had people do things like this... developers start to kick back. They end up with a very chatty API and lots more code than they need to perform a simple action. The audience of an API remains the developer and keeping to a purity line just causes most developers pain (not all devs, some prefer purity).
Going back to the example in the linked cookbook, they had a bank and a deposit resource. You can reasonably expect that the API permits a new deposit, permits fetching information about an existing deposit, but in the case of a bank account won't allow you to edit or delete a deposit once it's been made... you need to make a new deposit to fix that.
So reasonably links should have been returned that pretty much said:
<link rel="deposit" methods="HEAD,GET,POST" href="/account/12345/deposits" />
And for a specific deposit: <link rel="self" methods="HEAD,GET,PUT" href="/account/12345/deposits/17" />
DELETE wasn't allowed for either, and PUT and POST depended on the end point.I shouldn't do that though, as 'methods' is a gibberish attribute I just made up and no client will know what to do with that.
I should use OPTIONS, but then if you follow that through why include a 'type' attribute as you could've got that info from OPTIONS? (The entity-body of the OPTIONS response could give you the type information and even describe the schema of the resource).
I guess where I'm at is that once you start printing links according to what the user (not client) of the API can or cannot do, you're effectively echoing permissions through the use of links. And that permissions go beyond which resources can be touched, and into what actions you can perform on those resources. Meaning that to express this properly we start needing to be able to communicate which verbs are good for this user, for a given resource. Requiring a second HTTP request for everything to check these permissions seems a bit crazy in practise (very chatty APIs) though great in theory (conformance to every spec there is).
And what we're doing at the moment is opaque as the information on those interactions that the user can perform is held in different places, in the link, in OPTIONS, and additionally in specs. To a developer implementing against this, they're not given an easy way to just answer the question "What can I do right now?"... we give them half the information they need and leave it as a job for the developer to figure out the rest.
BTW: Respect for your book, thanks for replying as I'd love to hear your views on the above.
<3. Trying to get more good stuff out there...
> Then the developer has to hope that the API implementor also knew about that spec and did the right thing.
This is why relations that aren't specified in your media type or in the IANA list are supposed to be URIs. If they're not, they're breaking the Web Linking spec.
> But the missing part of that story is "and to make a call to this particular end point you need to use one of these methods: HEAD, or GET".
There is no reason that text cannot be in every link relation. It's just not. When defining your own, absolutely add that in.
> And that permissions go beyond which resources can be touched, and into what actions you can perform on those resources.
This is certainly a good insight.
> Requiring a second HTTP request for everything to check these permissions seems a bit crazy in practise
I agree 100%.
Ack! My pagination system is wrong! I'm going to try fixing that tommorow! Thanks!
"collectionName": {
"total": 861,
"limit": 10,
"offset": 100,
"pages": 87,
"links": [
{"rel": "first", "href": "/api/things?limit=10"},
{"rel": "prev", "href": "/api/things?limit=10&offset=90"},
{"rel": "self", "href": "/api/things?limit=10&offset=100"},
{"rel": "next", "href": "/api/things?limit=10&offset=110"},
{"rel": "last", "href": "/api/things?limit=10&offset=860"},
],
"items": [
...
]
}
It generically returns what was asked for. The items array... [{},{}]
I have seen the "total records" value return in the header with "206 Partial Content" response, but that still leaves the very important "links" information. Right now I am putting a "link-href" value for each object returned, but I don't know where to put the collections link information.Any thoughts are welcome and greatly appreciated!
EDIT: Adding to this comment because pfraze's comment below (https://news.ycombinator.com/item?id=5471523).
It seems as though the Link Header item could be the appropriate place for this: http://www.w3.org/wiki/LinkHeader
Where does the meta information about the resource belong? In the object/resource returned OR as meta information in the header (with other HTTP information). Is this a needless hindress for
There's a purity vs pragmatism argument, and I have tried the purity approach several times without success. This time I'm just going for pragmatism... giving developers an easy to use interface that is predictable and consistent.
I just didn't find it constructive to continue pursuing purity when the developers trying to implement the APIs just wanted to get the job done as simply as possible, with all the info at hand, and for them to return to what they were trying to do (usually solve some problem for a user).
But I only use //example.com/ if an API is truly available on both http and https (very unlikely, nearly all APIs should be on https only if they use some form of access token in the querystring for auth), and I only use https://example.com/ if the end point exists on some other domain.
As there is no technical difference I opt to save some bytes in the bandwidth.
Maybe it wasn't the author that submitted this? Or is the author asking for some help?
GitHub: https://github.com/restcookbook/restcookbook Authors/Contributors: https://github.com/restcookbook/restcookbook/contributors
1 - follow the AWS API models, with a signed request using a private secret known only to the user and the server-side. You can see the S3 docs on RESTful auth using this approach. Also seems to recommend doing this over SSL.
2 - use SSL and send a userid/passwd or authentication key on each request.
In general, cookies are regarded as one of those "makes it not restful" type things.
I'd love to hear from HN'ers on how they handle RESTful authentication, particularly for projects where they are providing an API that is primarily consumed by a web app or other tool they implemented for users and have used RESTful api design as a design viewpoint.
Here's what I do for that case: Create a /sessions endpoint and POST to that when you login. The session resource will include a token of some sort identifying the session securely. The client can then authenticate to the API by passing this token via an HTTP header.
There is no predefined way to deal with link discovery in JSON.
:(
How is linking now special in the sense that it can't be considered bona-fide JSON, if it is parsed by a JSON parser without error just like a shopping cart it.
Yeah I am not sure how important it is to say that I am serving applicaiton/foo+json or just application/json? I have been doing just application/json lately, I may very well be wrong about it.
The only advantage I can see to using custom mediatypes is if you want to do versioning (+v1, +v2, etc).
> How is linking now special in the sense that it can't be considered bona-fide JSON, if it is parsed by a JSON parser without error just like a shopping cart it.
Here's the definition of application/json: http://www.ietf.org/rfc/rfc4627.txt If my client accepts application/json, this is the only information that I can rely on for processing. Show me where links are defined in that spec.
They're not! Basically, we're having an implicit vs explicit argument: you're arguing that 'everyone should just know it's a convention I'm using, who cares?' and I'm saying 'By being explicit that you're using a convention, communication improves and everyone is more happy.'
This page doesn't even try to explain why HATEOAS is important. And I'm not sure that it is.
This example uses some arbitrary xml to describe the API, but if no client can consistently read and understand it, what purpose does it serve?
I believe, examples must be as real-world as possible. Say, the PATCH example must add at least an Content-Type header (and probably use RFC5261), otherwise an important aspect's lost.