Best practices for REST API design (2020)
stackoverflow.blog
stackoverflow.blog
EDIT: To be clear, I don't particularly care if an API is "RESTful" or not, as long as it's well-designed and documented, but I think there are interesting/useful ideas in REST that are lost when it's conflated with JSON-over-HTTP.
Please, just don't do it. Yes, I read the dissertation. It's just not a good idea, it has never given me any practical use whatsoever and has always made dealing with the API more annoying.
> Most of it is just standard JSON-over-HTTP stuff that's implemented in a variety of frameworks and libraries.
Yes, and we call that REST or RESTful anyway. We'll keep doing it that way. You're going to have to accept it.
The other option is munging URLs on the client side, which can be tedious.
I'm genuinely curious as to how including links in the response could make an API harder or more annoying to use.
Never overcomplicate things. Never add metadata if you don't have to. Just K.I.S.S
Somehow everyone rushed to implement GraphQL schemas, and no one rushed to implement JSON Schemas for example. And JSON Schemas with the right tools are a much more powerful instrument.
Same goes for linking, and the rest of HATEOAS.
At a company I worked we had an internal tool built for most of this stuff which took a lot of problems with HATEOAS away. Unfortunately, we never opensourced it :(
REST isn't really designed for the problem of implementing a single, centrally controlled API, and is basically overkill for that. Its made for a decentralized, heterogenous network of resources controlled by different parties, which can have separately evolving relationships. Like, say, the WWW.
If you have sufficient tooling around it to consume that kind of heterogenous resource network, it's trivial to apply it to a centralized service, but we don't really have the kind of client tooling for applications other than UIs that makes that nice widely available yet, and everyone wanting walled silos instead of free remixing means that REST is probably an antifeature for lots of API providers even if the tooling was 100% there.
If you believe REST is not a good idea then you never had to deal with versioning and struggling to keep clients and servers you don't own to play nice. Asserting that something like REST has no practical purpose is asserting that your experience is slim to none in this domain so that you are not mindful of the most basic challenges of getting servers and clients that evolve independently to continue to interoperate with minimal development effort. REST is content discovery and allowing clients to transparently adapt to breaking changes in the server. How does anyone with any relevant experience miss the point of that?
Furthermore, your assertion makes no sense at all. The main property, and the whole point, of REST is HATEOAS. It makes absolutely no sense to claim an API is REST if it misses the single most important design element that's behind REST. This is not pedantry or nit-picking: it's the whole difference between plain old RPC and REST. If you want to design an API that provides fixed endpoints to be called,and isn't discoverable or navigable, then just call the spade a spade: RPC over HTTP. Otherwise why do you feel the need to claim you design an API around a design principle you don't use and even criticize?
Sincere question: how does HATEOAS helps with this?
The only thing that HATEOAS helps with, is when the URL pointing to an entity changes, which is the most simplest change someone can make to its API.
But if an API ever changes the relationships between entities (new entities, 1-1 relationship changed to a 1-m relationship, ...) then HATEOAS won't help at all. You'll still need to change the logic of your client to take into account the modifications.
I've always felt like nobody bothered supporting HATEOAS because it's practically useless. The only use-case where it can be used is if you want your API to be crawled by a search engine. Then the search engine could crawl between entities thanks to the links you provide.
The whole point of HATEOAS is to address this. This is pretty much why it was developed and presented. Why on earth is anyone discussing REST not picking up on it's reason to exist?
> The only thing that HATEOAS helps with, is when the URL pointing to an entity changes, which is the most simplest change someone can make to its API.
It really isn't. At all. You're somehow missing the whole point of a resource-based architecture, let alone REST.
The whole point is that there are no fixed paths. At all. All there is is resources, whose representation might vary in shape and form, which are made to be discoverable by providing semantic descriptions of said resources. With REST you provide resources and metadata that points to representations of those resources. With REST your clients do not care about endpoints other than a root one, and resources. The resto of the process consists of tracking resource representations that you want, and you do not care where they are. At all.
This sort of profound misconception of what REST is supposed to be is the reason why somehow some people believe that it's reasonable to slap the REST label on APIs just because they churn out JSON. You can't. REST is based on design constraints that address a problem which these APIs fail to address at a fundamental level.
HATEOAS basically involves two main things:
1. Resource representation specifications are communicated in the communication channel (in HTTP, this can be a combination of MIME type in headers are more specific specification in the document itself for responses, and accepts header in requests.)
2. Related resources are identified by URL.
> The only thing that HATEOAS helps with, is when the URL pointing to an entity changes
False, though #2 helps with that.
> But if an API ever changes the relationships between entities (new entities, 1-1 relationship changed to a 1-m relationship, ...) then HATEOAS won't help at all.
No, that's what #1 helps with, as it provides the means where you can:
(1) identify the resource representations available,
(2) identify if you have a resource-consumption/creation-client available that satisfies the endpoint requirements, and
(3) select the correct client implementation.
> You'll still need to change the logic of your client to take into account the modifications.
Yes, someone needs to implement a new resource client if a new resource structure, or representation of the same structure, is developed.
And if you are doing a single, centralized API, where no one else will be hosting something using the same resources, REST doesn't provide much. Where REST shines (which is why it is abundantly used on the web) is where large numbers of different parties will be hosting services using a potentially shared set of resource representations, where each hosts may evolve the particular set in use on their own schedule, and perhaps even using different network protocols, and you don't want to have bespoke clients for each service, but instead protocol clients for each network protocol and resource clients for each resource representation, and in-band signalling of not only where to get each resource once you get to an entry point, but also have the network protocol client you need and the resource client you need to successfully carry out each transaction.
I'm not fine with making an API use HATEOAS, especially not for the sake of being pedantically accurate about the word REST.
As for your defense of HATEOAS, you're not doing a good job. If it's such a good idea, why is pretty much nobody doing it as intended? How exactly does it help with versioning? Can you describe a real use case?
The way I see it, if you want to break your API, slap /v2 in front of it. Otherwise, I strongly suggest that you do not break your API. Do not move stuff around pointlessly. I guarantee you that even if you followed HATEOAS best practices to the letter, your clients are still going to break. You don't want that.
And that's perfectly fine, but don't call what you're doing REST because it really isn't.
There are many other industry buzzwords like "object oriented" that don't match up with what the people who coined the term had in mind. It's a nuisance, but nothing to get worked up over.
So what shall we call REST without HATEOAS[1]? "RESTless"?
With only links, you don't know semantics and inputs. The benefits are next to 0, and effects on the maintenance are non trivial.
- How to handle date and time incl. time zones - JSON has no data type for that - Handling of numbers (JSON only has double, which does not fit most cases) - Defined and parseable error responses (rfc 7807 plus extra fields for details) - Localization - do you send translated texts or just error codes? - How to handle updates? Overwrite every field? How to update only some fields for a resource? Optimistic locking? - How to handle 1+n problems for reading? ...
Send numbers as strings. you should be treating this as hostile in your backend anyways, and checking it.
Send both an error code, and a string in simple english, or whatever your most common developer language is.
If you care about only updating certain fields, track changes and only send those fields to the backend.
These arn't really that hard.
In other news we’re suspecting transactions are on horizon. Half the backend team are in denial, the other half are busy trying to figure out the most pleasant place to install decidedly “unRESTy URLs” for them. Somebody already set themselves on fire over PUT/POST for that.
With offset, ok. As UTC alone, it doesn't work if you need to know the local time when something occurred, like medication administration, especially when the things may occur across time changes like DST.
The endpoint of the viewer may not be in the same local time.
If you go this route you may not care about location beyond this one data point, but have to store another location field each time.
Why? Why not send a number as number (double) and treat it as hostile in the backend? I dislike sending numbers as string because I think the different data types exist for a reason.
Same thing with serial numbers, VIN numbers, building floor numbers, phone numbers, etc. All of those should always be strings, because performing "math" on them wouldn't yield meaningful results.
Why should i send a dimension as a string? How am I going to calculate the volume from 3 strings?
You shouldn't because a dimension is a number. Something 4 feet long is twice as long as something 2 feet long.
So prices are numbers (if supplied in cents, the smallest unit). There is length/depth/height/weight/... A lot of things are actual numbers. Why supply them as string?
I agree partially with you but think ZIP are a very bad example because a zip itself may also contain letters.
Either that or use a different serialization method to communicate.
No one else seems to be complaining about this issue. If JSON doesn't fit your needs, use something else.
what's an ISO date/time?
> JSON.stringify(new Date())
'"2021-02-22T20:34:53.686Z"'That said, "time zone designator" is the formal name of the UTC offset in ISO 8601, and it's reasonable to refer to it as such whenever the distinction is clear.
Handling Post and Patch differently can at least give a general guideline for "overwriting vs keeping fields on update."
The localisation part is trivial if you consider that APIs are to be consumed by machines - much harder, if you assume they're used for GUIs.
Error handling is easy to define right, but so tedious to implement....
1+n can be improved by "guessing" which nested ressources are going to be fetched, and include them "by default."
But you're right that those are decisions that need to me made, and can quickly be agonised on without adding much value....
This distinction is important because you can't serialize infinities or NaN, and there's no guarantee the JSON number can be accurately represented as a double. JS likes to pretend that JSON number is interchangeable with Number and this can result in some fun situations when your Infinity becomes null
I guess the point is that JSON has about as much to do with JS as JavaScript has to do with Java.
Slicing by ID is how Github does commit history and it drives me nuts that I can't jump several pages, for example, to see when the first commit was, or to guess whereabouts some commit is given a known time range. IDs make it impossible to do anything but iterate step by step.
I much prefer an interface that exposes:
a) how many items exist in total
b) which offset it's starting from
c) how many items in the slice
Then if I want a slice from 300-8000th items, I can type exactly that in the URL. Yes, I understand this will render a huge page, just let me do it the one time so I won't be spamming your server with requests over the next hour trying to find something while fighting against bad UX.
Fair point. I agree, that's a nice side effect. But what if you're looking at page 1, and an item is removed from that page? Then you'll never see the first item at page 2, because it's now the last on page 1.
> Then if I want a slice from 300-8000th items, I can type exactly that in the URL.
That's a nice feature. But it can put a lot of load on your backend if you paginate over 10 of thousands of items.
How common are concurrent removals in practice though? I can't think of a single instance off the top of my head where this is problematic and you can't just go back to the previous page if you're really paranoid or confused that something is "missing"
> it can put a lot of load on your backend if you paginate over 10 of thousands of items
Look at it in aggregate. Fiddling w/ URL params to paginate (instead of clicking pagination links) is a power user move. If I want the date of first commit in a repo, I'll only look at the last page (vs paging through the entire history). For guessing, I can click a page, and if I went too far, binary search from there (again, vs linear search). Etc.
Even for the most degenerate use case (e.g. some jerk trying to crawl over the entire dataset), the load is smaller with a single request than the overhead of multiple requests. Paginating is not an appropriate mitigation strategy against this type of traffic, and you arguably can implement detection/caching/blocking mechanisms much more easily for naive huge queries than if you need to differentiate regular traffic from bots.
The whole issue is that you're never going to know about it. Sure, you can write some convoluted automated process to double-check previous page(s) but most people aren't going to do that.
Another thing to consider when paginating via id cursors: if the deleted item is the one your cursor is sitting on, then you no longer have a frame of reference at all. For example, what happens to commit history pagination in github's implementation if I rewrite git history? Chibicc for example deliberately rewrites its history for didactic purposes.
It's difficult to do in most HTTP deployments where we try to limit HTTP response duration.
> if the deleted item is the one your cursor is sitting on, then you no longer have a frame of reference at all
Not a problem. If your order your results by timestamp and ID, then you provide the timestamp and ID of the last item returned, and take anything after. Doesn't matter if the item is still there or not.
To prevent that, usually maximum page size is enforced on the server anyway and the client is informed about the actual page size in the metadata in the reply.
It's a trade-off.
Thus using page numbers is probably a pretty poor proxy for what you're actually trying to do when you say "getting to arbitrary pages." Presumably you are wanting to skip to a specific place in the list, perhaps specified as a percentage ("take me halfway through the list") or as some predicate on the data ("take me to items from 2 weeks ago"). APIs should provide ways of expressing these specific places in a list, instead of requiring you to either guess page numbers or do extra work to calculate them.
You'd still present the user with a set of results (a "page") and would let them seek forwards/backwards to the adjacent subsets of results.
AFAICT cursors can't support this.
And even if you go halfways into the list you can't display all subsequent items, so you'd still have to paginate in some way.
The page metaphor is there for a good reason.
The other concept is using an actual page number to make requests, e.g. requesting {page: 1} and then subsequently requesting {page: 2}. This concept is the one I was claiming is less desirable than some alternatives.
As for cursors, I don't see any reason why you couldn't make requests like {listPosition: "50%"} or {createdBefore: "2020-02-15"} and then still use cursors in the response to request the previous or next page. (Those two examples probably aren't actually good API naming conventions, but it should demonstrate the idea.)
1. APIs consumed by other backends lets call them API2B
2. APIs consumed by frontends lets call them API2C
In this case cursors are better for API2B but not for API2C as in case of API2C most users expect to be able to jump directly to a specific page.
At least when I am designing an API i take different decisions based on this split. For example in case of API2C I always want to see FE design even if I work on backend.
Which ones?
for example GET /customer/1/orders/
response:
{'orders': [order1, order2, order3],
'navigation': {
'firstpage': '/get/customer/1/orders/',
'nextpage': '/get/customer/1/orders/?query=orderId^GT4',
'lastpage': '/get/customer/1/orders/?query=orderId^GT990',
'totalrows': 1000
}
GT means greater than
Edit: obviously the proper impl totally depends on your app. If the result set is enormous, caching doesn't make sense. If it's rapidly changing, pagination probably doesn't make sense. etc, etc.
That said, using “page of results preceding <first id if following page>” and “page of results following <last id of preceding page>” reduces obvious pagination artifacts compared to <page number n>. Whether it's better to beat consumers over the head with inconsistency due to concurrent changes probably depends on the application or provide a nearer illusion of consistency probably depends on application domain and use case.
If we don't care about added items (they can be deduped client-side), and only care about removed items, maybe the backend can maintain a tombstone timestamp on each deleted item, instead of deleting them, and then the client can provide a "snapshot" timestamp in the query that can be compared with the tombstone timestamp.
A query ID is super easy to implement, and yes, for any application you run you have to be aware of the resource requirements. Tune your eviction policy, and this is totally feasible.
Deduping and tombstones is messier to implement IMO, and hitting a cached result for the next page (e.g. a redis LIST) is probably less expensive than a "query" (whatever that means for your backend).
Padding the offset would solve for the problem mentioned where deleting an item would mean some non deleted items are not included in the paging results because they got moved up a page after that page was requested, but before the next page was requested. For example If I request 100 items at a time but set my offset to be 90 more than what i have received so far, i can expect my response to have duplicates, if it does not then i know my offset was not padded enough and i can request from a different offset. Of course you would adjust the numbers based on knowledge of the data.
Edit: If you used the ID of the last item instead of an offset, then you could get errors if your last item is in fact the one that was deleted.
> Edit: If you used the ID of the last item instead of an offset, then you could get errors if your last item is in fact the one that was deleted.
If the items are sorted by ID, then we use the ID of the last item x. But for example if they are sorted by timestamp, then we use the timestamp of the last item (and maybe the ID to break ties). Then it doesn't matter if the last item is still there or not. We are only considering values before or after depending on the sort order.
for instance your comment is 26227524
and the parent post is 26225373
edit: hacker news pagination is not high priority apparently since they just use the easiest way with p= some number.
Folks that aren't aware of Webmachine should take a look:
https://github.com/webmachine/webmachine
The 'Accept' header should determine the response type, but content negotiation is something that few bother to implement. Webmachine does that for you, among other things like choosing the correct status code for your response.
Also, shameless plug for my OCaml port:
Great buzzwords, I have no idea what this project actually does.
The only advantage I see is that your clients can choose the format they want to work with, but since the serialization format is just a message format that has no impact on the code, and that most of them have a bijective transformation between them, I don't see the point.
If it feels like a lot of work (Or the added complexity of a moving part, if using the tool you linked) for next to no impact.
if you do it right, using the standard html content type returns a human browseable representation of your api with forms for posts and whatnot.
Also, the HTTP decision diagram: https://github.com/for-GET/http-decision-diagram/tree/master...
cat -> cats,
dog -> dogs,
child -> children,
person -> people,
etc.
It makes my code feel inconsistent too. Sometimes I use more specific typing to resolve this, e.g. PersonList , but I'm not sure that's any better.
In this scenario, consistency is more useful than grammatical correctness [1]. And the simpler pluralization logic makes it easier for non-English speakers to work with the codebase.
[1] Related example: https://en.wikipedia.org/wiki/HTTP_referer#Etymology
Things are object_detail (singular) and object_list (plural).
It just assumes 'append s' for plural displays, unless you override. I like the _list suffix instead of trying to pluralize
- Accept and respond with JSON: No, why? what if I want to work with XML? or format X or Y?
- Use nouns instead of verbs in endpoint paths: I don't care about that
- Name collections with plural nouns: I don't care about that
- Nesting resources for hierarchical objects: No, why?
- Handle errors gracefully and return standard error codes: nothing to do with REST specifically.
- Allow filtering, sorting, and pagination: yes, good luck coming up with the right query structure and sticking to it.
- Maintain Good Security Practices: nothing to do with REST specifically.
- Versioning our APIs: And here we are. How many HATEOAS people claimed it was blasphemy and it missed the point of REST?
None of these are an issue with GraphQL, just like they weren't an issue with SOAP. You get a schema both the client and the server must agree with, end of story.
As I said to a dev once, if your 'best practices' can't be automated with a CI tool, then you need to worry about creating that tool first...
You want to know how to halve your performance and responses/s on your service? add graphql.
All things that GraphQL claims to do can be implemented in RESTful services easily.
If you want specificity in your query fetching, just add query params or put them in the request body
If you want schema validations, there are many libraries that help you with that.
And if you want data from multiple resources from different endpoints, what exactly is stopping you from implementing that in REST?
Also, GraphQL has a steeper learning curve as opposed to REST.
never understood graphql. To me its just an abstract layer between the front end and back end, adding to the already complex stack.
From a backend perspective, it's about 100x more work than dumb RPC (which I prefer over "REST"). I have spent so much time becoming aware of very "interesting" decisions. For example, gqlgen for Go generates code that fetches every element of a slice in parallel. We used to create one database transaction per HTTP request, but you can't do this with GraphQL, because it fetches each row of the database in a separate goroutine, and that is not something you can do with database transactions. It's also exceedingly inefficient.
All in all, it's clear to me that GraphQL is built with the mindset that you are reaching out to an external service to fetch every piece of data. That makes a lot of sense in certain use cases, but makes less sense if you are just building a CRUD app that talks to a single database.
I am hoping that people get bored with this and make gRPC/Web not require a terabyte of Javascript to be sent to the client. RPCs are so easy. Call a function with an argument. Receive a return value and status code. Easy. Boring. It's perfect.
With GraphQL, you're just pushing N+1 calls and figuring out the mental contortion needed to support all the graphy-ness from non-graph structures. Most that say GraphQL is the bees-knees must be an FE developer. There's nothing wrong with that, but GraphQL is just NOT as cracked up to be.
Eg: Hasura.
https://hasura.io/blog/architecture-of-a-high-performance-gr...
How do you easily write an endpoint that can either return a comment, a comment with its children, a comment with its corresponding story, with a strongly typed schema (ie. no optional children or story whatever the query is)?
it's not hard.
In my specific case, using Flask-Marshmallow i can do this
```
some_resource = Resource.find_by_id(resource_id)
arg = request.args.get('type_of_resource')
if arg == 'with_commments':
return full_resource_schema.dump(some_resource), 200
if arg == 'no_comments':
return no_comments_schema.dump(some_resource), 200
```there i just implemented graphql in 6 lines of code
Dare i even say switching to graphql and onboarding developers to graphql is a much more resource intensive process than adding another endpoint (or a couple).
As a BE/remote I basically taught myself most of elixir's graphQL framework in a day (wait, that was yesterday), while under feature pressure (due this afternoon), mostly by poking around and doing a bit of TDD. There are parts that I really hate about graphQL, but overall I consider it a win.
Instead of doing a simple GET /a/b/1/c and presenting that data structure, they now need to define a query, think about what resources to pull into that query etc. If the query ends up being slow, they have to understand why and work on more complex changes than simply asking the BE team for a smaller response with a few query params.
We quickly realized that expecting them to learn our data model and how to use it efficiently would be much more complicated than exposing every plausible use-case explicitly. We could do this on the "API front-end" by building a set of high-level utilities that would embed the GraphQL queries, but that would essentially double much of the work being done in the front-end (and more than double if some customers want to use Python scripting while others want JS and others want TCL or Perl).
So, we decided that the best place to expose those high-level abstractions is exactly the REST API, where it is maximally re-usable.
Most API teams will give you all the resources that you need.
It is most likely the case that the FE team will ask not for more, but rather, for less.
This was the exact problem why GraphQl was created in the first place - to account for Facebook's mobile platform, because they were receiving way too much data that needed to be trimmed because they didn't want to be sending like 500kb of json over mobile networks.
Competent API teams will give you all the resources that you need to be productive as a FE engineer, and if it is the case that your data needs some trimming, then go ahead and ask your decoupled backend team, while you can go ahead and keep working asynchronously while they get that done for you.
Not Fatal as you think it is.
Graphql works the same way if i'm not mistaken.
and if its the case that you don't want the comments at all, then in the conditionals write your sql statements
if you want comments
select * from resource where id=some resource id
if you don't
select without comments from resource where id=some resource id
there's no need to get into specifics, the point is graphql is not bringing anything novel. everything it claims to do can be implemented in RESt if you're willing to implement it.
With a REST API, the “resource” type will always have an optional “comments” field, whatever the value of arg is.
With a GraphQL API you’ll have the right type depending on your query.
In general, the tooling/spec has thousands of open bugs, isnt going anywhere, moves very slowly and is missing fundamental features.
https://github.com/OAI/OpenAPI-Specification/issues/1998 heres the ticket where it will inevitably be ignored and die. But even if it was supported in the spec, the open source generators wont support it anyways.
Said generators also uses "logic less" templates to generate something that very much needs logic. As a FE dev, I love nothing more than to fork a an archaic Java based generator to get the client gen working...not.
So if you wanted comments you would send a request with a header like "Accept: application/resource_with_comments", and you'd send a Header like "Accept: application/resource_without_comments".
The thing is, few people want this level of control in practice. This would generally lead to a huge proliferation of types for very little practical benefit. Much simpler to have a set of well-defined types with optional fields than to define a new type for each combination of fields used today.
- some fields of the resource - some fields of the comments - some fields of the user that made the comments
API consumers definitely aren't the ones who should be thinking about something like joining behavior, since GraphQL doesn't give nearly enough control to achieve performant joining.
Basic filtering can also be easily achieved with ad-hoc methods, such as query params. No need for a full-blown GraphQL language server in front of all your services.
At least, this has been our experience with B2B software sales.
Of course, YMMV depending on market etc.
Your API should match semantic actions on your client like the UI and it needs to do this in such a way that it happens in one database transaction.
With REST, new UI requirement that needs to save 2 objects together across multiple entity types...? GG.
The TL;DR of that blog is this. There isn't much value that GQL can provide over REST + OpenAPI. GQL is still in its infancy so APM is easier with REST. The biggest argument for GQL is returning the right amount of data but GQL is slightly slower than REST because the server is having to parse what is essentially a mini program with every request.
In my experence therefore, the tooing (code gen, doc gen) etc around OpenAPI is piss poor in comparison to gql. Just look at GraphIQL compared to postman.
GraphQL has first class support for this scenario.
The whole point of an API is to decouple. GraphQL does the exact opposite.
...may not be the advertising you think it is.
I believe that treating RESTful APIs as a smart data storage that can do advanced querying is a mistake, or rather REST is not the proper solution for this use case, and you did well to move on from it. Personally I've never failed to use REST successfully as long as I was willing to do proper caching, sorting and filtering in the API consumer. REST is plenty if you need presenting collections of things with key-set pagination and proper caching headers.
In the days of HTTP2+, it's easy to fire 100 concurrent requests to a key-set RESTFUL API and load all the data that you could ever need. :)
But I have missed feelings about graphql. I wish they didn’t invent a new language, it was just json. I’ve encountered so many little bugs because graphql parsing is different between different servers.
Ideas of graphql are great, implementation seems over complicated.
REST is much easier to grok.
It's certainly one of the main sales points for graphql. On the flip side, I've never been frustrated by getting too many fields back from an API. I suppose if I was developing exclusively in extremely bandwidth limited contexts where getting back only 2 fields rather than 50 actually made a difference, I might care. It just seems like such a non-issue to solve/talk about in almost every other context.
This works way nicer with caches and APIs where the URI is a first class citizen.
A pretty common use-case I have is needing to support these three things :
- A “big object” list screen (where retrieving the whole objects would make the query return megabytes off data)
- A “big object” details screen (where I need the full object)
- Programmatically getting many big objects
With GraphQL it involves writing one (or two) straightforward queries, while with REST it would require more thought or code.
Aren't you just pushing the work to the back end? The GraphQL resolver is a new layer of complexity while with REST is more straightforward.
If there is some sequence that is used a lot and you've determined it to be a bottleneck, you can write a custom endpoint for it.
I could fetch the objects in bulk request.. except they didn't have the field i needed when fetched in bulk. So i had to do bulk fetch, where each object was pretty big and then fetch them again individually.
If that service supported graphql i could easily just grab what i need, without hammering the poor already over-utilized server.
This kind of issues prop up constantly especially if you use on-perm hosting.
You could throw static json blobs on S3 as a rest API but selections would not be supported.
Amazon S3 Select works on objects stored in CSV, JSON, or Apache Parquet format. It also works with objects that are compressed with GZIP or BZIP2 (for CSV and JSON objects only), and server-side encrypted objects. You can specify the format of the results as either CSV or JSON, and you can determine how the records in the result are delimited.
You pass SQL expressions to Amazon S3 in the request. Amazon S3 Select supports a subset of SQL. For more information about the SQL elements that are supported by Amazon S3 Select, see SQL reference for Amazon S3 Select and S3 Glacier Select.
You can perform SQL queries using AWS SDKs, the SELECT Object Content REST API, the AWS Command Line Interface (AWS CLI), or the Amazon S3 console. The Amazon S3 console limits the amount of data returned to 40 MB. To retrieve more data, use the AWS CLI or the API.
https://aws.amazon.com/blogs/aws/s3-glacier-select/
https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-gla...
The "; charset=utf-8" part is not necessary and is ignored. See https://tools.ietf.org/html/rfc8259 which defines the application/json media type: "No "charset" parameter is defined for this registration. Adding one really has no effect on compliant recipients."
https://github.com/microsoft/api-guidelines/blob/vNext/Guide...
I was always under the impression that 409 would be the "correct" code here, while you'd use 400 if the API user didn't supply an email at all, for example.
No user cares whether an error response came back with a 400 or a 409 status, or what those codes even mean. It's madness that as a profession we spend so much of our employers' time and money on trivial things that deliver no value whatsoever. Response codes, HTTP verbs and pretty hierarchical URLs are all meaningless. I've lost count of the number of meetings I've been in where there have been pointless, timewasting debates about these things. Should we expose separate /foos and /bars endpoints? Should it be /foos/{foo_id}/bars/{bar_id}? Or maybe both? Do we need a POST, a PUT or a PATCH? Etc, ad nauseam.
I'm generally cautious when it comes to embracing new technologies, but I've wholeheartedly embraced GraphQL simply because it has removed countless hours of unproductive bikeshedding from my professional life.
It is a bad engineering idea to reuse HTTP codes that can be thrown by multiple intermediaries confusing the client code.
That said, some of the practices (like being overly clever with status codes) are pointless while others are actually helpful design patterns.
Using nouns over verbs helps you fully think through the process of mutating state in a RESTful/stateless way. The HTTP verbs all have a unique purpose you should understand.
Does it always matter what size of hammer you use? No. That doesn't mean its smart to use a screwdriver as a chisel.
Unless your product is itself an API, your users shouldn't be exposed to its status codes. Sorry for dispensing trite advice, but you should really have a nice client between API and users, that can translate the difference between 400 and 409 statuses in a comprehensible and friendly way. :)
In my experience, in any situation where you produce an API where you want the client to pay close attention to the actual error that occurred, the list of official HTTP response codes is generally woefully inadequate for the task and you're going to want to work with something much more specific.
If my payment processor is telling me they're declining a transaction because the customer failed a fraud check, I'd rather they didn't communicate that to me with some random 4xx code one of their devs found on a Wikipedia article and tried to bend into shape. Rather, I'd hope the importance of the situation would lead them to compile a table of custom error codes that I can implement properly in my program's logic.
Won't you just send up in those huge discussions you waste so much time on?
That doesn't mean I wouldn't also use the best practices codes for authentication vs authorization etc. I say this loosely though because best practices are really just that. It's more important to maintain a level of consistency within your team or organization so you're all speaking the same language. It's nice to have a starting point like rest APIs (or something that resembles them) so you're not writing a completely new rulebook on standard stuff that other people have already figured out, you can focus more on the problems your team is trying to solve. Why not use a 409 when there's a conflict? Someone already solved that problem. You can still return whatever pretty error message you want.
4xx codes are processing errors, or application-level refusals.
I end up borrowing 422 from WebDAV for a generic "Unprocessable Entity", with an error code (for API interpretation/reference lookup) and description (for user display, when appropriate) in the response body.
By ignoring existing standards and re-inventing the wheel with custom errors the previous development team made the system harder to maintain & harder to onboard.
maybe I'll investigate graphql a bit more, I've just heard it was really complicated to implement the resolvers for the queries so people weren't using them yet.
The ambiguity is probably why most people just use 400 and call it a day.
In the given example we work with an Article domain. Where there are Comments.
Fine and totally fair.
But how would you design your object graph from the bottom up?
Articles and Comments is a fairly simple example.
Imagine a Transaction in bank domain context.
Would you have an endpoint like so?
transactions/:id/acoount/:id
Or perhaps
transactions/:id account/:id/transactions
Whatever you pick you must find a way for transactions and accounts to work together on a design level, which will impact how you potentially serve services and how you deploy them.
Is an account and a transaction part of the same domain and service?
It's hard to design proper REST apis because although it's easy to dish out cool looking routes, it might be a lot harder to actually design your systems based on those routes.
You don’t need hierarchy in your URLs. If you want to load a transaction, then /transactions/123 is fine. If you want to load an account, then /accounts/456 is fine. You don’t need to put one under the other.
In a REST API, the primary key for a resource is the URL. If you have the primary key, URL hierarchy is unnecessary. Just use whatever is clearest for a developer to identify at a glance for debugging purposes. It’s not required by the tech.
https://docs.microsoft.com/en-us/azure/architecture/best-pra...
This has massive implications in real-world systems. Rate limiting adds and order of magnitude complexity to a client system due to the next for delayed, conditional execution. So many APIs do rate limiting and don't bother to add information about why it is happening or for how long the limit applies. I've been working with a name-brand system that limits based both on client application and end user. So, effectively, there can be a 429 because the application as a whole is exhausted, or because a particular user is exhausted. There is no additional information, so the only way to know if the limit has lapsed is to try again. It make it incredibly cumbersome to create reliable consumer applications.
https://cloud.google.com/apis/design/resources
Also as one of the most common pushbacks people have with gRPC is that everyone generally expects JSON still in 2021 especially in the browser. Not sure if this is super well known or not but gRPC integrates well with things like envoy which allow you to write a single gRPC API and have it serve and respond to both native gRPC and standard REST / JSON calls.
What are the versioning problems with it? I was under the impression that it was actually designed from the start to be fairly flexible with regards to changing API methods etc
https://cacm.acm.org/magazines/2019/12/241052-api-practices-...
Many REST APIs have problems for the same reason SOAP had, and it involves lack of deliberate API design and abuse of "automagic code generation"
> 401 Unauthorized – This means the user isn’t not authorized to access a resource.
Surely that should say that the user isn't not unauthorised?
401 is unauthenticated (not logged in).
Very important difference for clients handling authentication state.
The only thing something like gRPC is missing that is very valuable from HTTP REST is having standard error codes corresponding to all the HTTP response codes (2xx, 4xx and 5xx codes)
It works for simple CRUD APIs, but it becomes too limiting once there are more than one way of updating something.
It ties the API design too closely to the data model and usually means that the client has to implement more business logic instead of handling that on the server.
Good API design should decouple business logic from the HTTP spec.
Instead of “doSomeAction on resource” it’s “POST/PUT /resource/actionRequest”.
The trick is stopping and trying to make everything work over CRUD of your main entity resources. HTTP resources don’t need to map to tables.
https://www.oreilly.com/library/view/rest-api-design/9781449...
https://www.vinaysahni.com/best-practices-for-a-pragmatic-re...
GET /resource/1
Accept-Version: ~1.0.2
...
Implementation-wise this ends up with better code (and less code) on the back-end and a reduced number of HTTP routes (which should represent the entities and be maintained).The pain of REST is that there's very little in terms of "correct" as we don't have a formal specification of the semantics. If you end up using /v1 it's not any more or less REST-ful.
Media-type versioning is useful for supporting different data types/representations of the same resource.
URL versioning is useful for versioning resources themselves and their behavior/business logic.
Good luck with scaling and automatic testing tho.
Although I would love for browsers to start natively supporting some sort of schema based, explicitly typed, binary format.
As usual only CRUD counts, no mention of operations, patch, links, aggregation of calls, ways to document API, token duration, rate limits etc.
Adequate for todo list maybe.
How do you distinguish between an invalid path (no end point) and a valid path requesting a resource that doesn’t exist?
I can't think for other types of not existing, or why you'd need to differentiate between them.
I call /customer/235235 and customer 235235 doesn't exist. That's a 404, resource not found.
But say I make an uncaught error with the path and call /cutsomer/235235. That also is a 404, resource not found.
It really depends on what you want the word "resource" to mean.
It gets a bit more complex if you have a simplistic website with an API service both on the same server. For your users, you want a nice 404 page for mistyped or dead links. But you don't want your API also triggering that same page.
Beyond actually splitting your main web content and API into two servers, I could see returning a 400, Bad Request response for the API to avoid triggering a 404 handler. After all if you give it a customer ID in the above case that doesn't exist, it is essentially a "bad input parameter" and the error message can indicate which parameter (the customer ID) and why it's a problem (there is no such customer).
To answer the specific question /cutsomer/123 should never be hit in a production application, unless your user actually inputs that, so treating it differently is just overhead that gives no benefit.
Also for the "both on the same server" scenario, you can make a pretty good assumption if a user agent is a browser or an API client, so you can present the right output to the right one. The accept header is the simplest method I can think of.
I once worked with an old API and when a pentest team ran a scanner against it they reported hundreds of backdoors and malicious code false positives because they were expecting a 400 error rather than an 200 with success=false response body
invalid path: 404
GET /api/vn/singular_name/ = list all
GET /api/vn/singular_name/?field=value = list all filter by query, if it gets more complex base64 encode the query and decode on server
GET /api/vn/singular_name/:id = get by id
POST /api/vn/singular_name/ = create
PUT /api/vn/singular_name/:id = replace
PATCH /api/vn/singular_name/:id = atomic update
DELETE /api/vn/singular_name/:id = duh
If models are related they receive a related_model_id field. This works even for "many to many".
In the case of Vue there's this excellent library called vuex-orm, which handles normalization of incoming data and ways to define relationships or all kinds and also queries with or without related models. It's so good that I'm abandoning Angular because it doesn't have something so sophisticated, despite it having the better component system and structure all together. Not having to subscribe and unsubscribe to data streams and deal with this torture instrument rxjs is a bonus. Because of vuex-orm I'm more productive than I've ever been with Angular.
This is sufficient and there is no need for GraphQL.
Real time updates are simple. Whenever something like INSERT, UPDATE, DELETE happens an event is sent to a notification library which sends those updates out via websocket, a client version of this handles updates of the cache (or if you like vuex state, previously ngrx-data cache). The notification library also handles scoped status updates. Imagine a multi-community site. You're currently in community A and don't care about what happens in community B, the client tells the server "I'm in A" the server notifies the client only about the latest happenings in scope community A. Once you navigate to community B you do a fetch from the server to get the latest version of the truth. But while you're there you're also telling the server "I'm in B, send me real time updates about B".
I'm not a fan of Redis. I like MongoDB and having a single database system for everything. Since I'm not into game development that's good enough. Postgres is nice for querying, I like SQL more than Mongo's way to query but when developing running migrations, database updates, it's a PITA, especially in Go which seems to lack tooling for this and the driver and "o"rm landscape is also a bit barren.