In typical HTTP REST services this transport/API error split makes it really easy to create client code which only needs to check two conditions - if the HTTP response code is 200 or not, and if the error value is set or not. You also don't have to shoehorn your error handling into the very limited set of errors provided by the HTTP protocol.
The other real-world advantage of this is that when you outgrow HTTP as the transport protocol for performance reasons this makes porting the API really easy to, e.g. protobuf RPC, or even raw TCP. The error is already defined as part of the API and you don't need to rewrite all your client code to deal with mapping multiple HTTP response codes to your new transport. It's good future proofing I've seen pay off in a at least a couple of real-world cases.
Bottom line - your server should always return HTTP status 200 and a separate API error response.
There's also a reasonable debate to have about whether non-error responses should also include an explicit "error" field with some default OK value. There may be good reasons to leave it out, e.g. if you want to save bandwidth, but I consider that a fairly insignificant point. For consistency my APIs always return a default OK error field on non-error responses - your mileage may vary.
From the client side, developers shouldn't need to worry about what piece of your stack was at fault. They want to know if it's your fault or theirs/their user's. And if it's their fault, how to prevent it or fix it. Both HTTP and gRPC are flexible enough to encode that information.
From the server side, it complicates monitoring in practice if errors are propagated encoded in the response body of a "successful" RPC. With reverse proxies and other things in place, the complication increases, and many people get unhappy.
W.r.t. migrating clients to other transports, the client library should hide that (we haven't done that always right in the past, but we do now).
Also note that not all APIs will be used be developers. Often it's someone less proficient with programming, and you can't rely on them to check the content of the return message. Help them help themselves by making curl (or whatever) bitch when there is an error.
It takes a bit more work to design an API that works over multiple transports, but a good framework, such as Servicestack.net if your in .Net land, mostly does it for you. By advocating 200 for all responses you're basically reverting to SOAP and WCF (Windows Communication Foundation).
Each to his own though, and for internal stuff, it might make a lot more sense.
>> Bottom line - your server should always return HTTP status 200 and a separate API error response.
That is a very absolute statement from a pragmatic guy, asking people to concede viewpoints other than their own :)
I'd be interested in learning how status codes other than 200 cause problems at scale though?
"Hi, this is the application and the object/document/whatever you are looking for was not found"
with
"Hi, this is the server and the api endpoint was not found"
An API I use returns a standard apache 404 error page when the item you are looking up doesn't exist. If the API endpoint was renamed my code that consumes it wouldn't have any idea anything was wrong.
I didn't mean to imply that HTTP status codes could stand alone as error messages, so for a 404 error I'd also respond with more data, to help the user identify the issue.
Huh? What networking model are you working on?
HTTP is application layer.
TCP is transport layer.
IP is internet layer.
Ethernet is network layer.
Disclaimer: I am one of the co-authors.
I don't recall that specific scenarios, but over the past year I learned this with the maps API. As long as you're hitting a real endpoint, then it will return 200 - even when there are errors and it should clearly return a matching http status code.
I suppose it so the same API could be implemented in any given protocol, however, I don't think that is useful in many cases.
Otherwise HTTP status codes for errors are used, and standardized internally.