Advice for Operating a Public-Facing API
jcs.org
jcs.org
A 200 may mean the payload has changed, and how would an app handle an arbitrary format change? This could well fall into the “don’t test for something you can’t handle”.
At a minimum it’ll ideally get the app to quiet down.
That premise is incorrect, as 4xx is client errors, as in the client formatted the request wrong or anything else, and therefore it couldn't be responded to. Rate-limiting is the client hitting the server too much, and it's up the client to handle this.
5xx is for server issues.
2xx should be success in any shape or form, so clearly 2xx shouldn't be used in this case either.
I agree that 429 should just be used for it's intended purpose here, handling rate-limiting requests/responses. You can still add a body if you want, with the current quotas.
It's not what "should" be used, it's what the author found to be effective.
But trying to handle clients who mishandle things like that is a fools errand. What client, in their right mind, would try to retry a request that is failing because of what the client is sending? In no case does that make sense, ever.
Similarly, should everything just be 200 then just in case clients mishandle redirect requests?
People will copy random snippets from SO and smack them with a hammer until they seem to work then move on to the next thing. I've seen some incredibly stupid code out there, code I can only assume the author either didn't understand or truly didn't give a fuck about. Probably both.
Sure, I agree a lot with this, but that doesn't mean you and me should also do idiotic things. Lets just return correct status codes and the ones who misuse it, will misuse it :)
This guy should not be designing APIs.
If you did that, and no error handling was implemented in the IoT device, it would merely result in a non-event that no one knows about.
In this specific instance, the author is doing the right thing. The developer will see truncated messages running through.
In the end it is a question on desired semantics. The semantics of the operation in question is obviously to truncate and process rather than fail.
(I actually agree with you, I think it's a bad rule of thumb. I prefer Fail Fast (and loudly)).
What makes you think that the API described in this post would be subject to those kinds of legal and/or regulatory requirements?
Could it be the case that the author knows prima facie that no such requirements exist for their service?
If I target use cases that prefer truncation of messages to errors, then surely it's acceptable for me to build it that way, right? I'm not saying it's a good idea in general. I'm saying it can be acceptable in the appropriate context.
If a client desires this behaviour, then add a `truncate=true` flag that has to be explicitly passed.
If that data doesn’t exist, public announcement that the API is breaking long messages without the new parameter.
If the API can’t be broken, version the API and break it in the new one and for new customers.
In all cases, disclosure in the documentation - If you know there’s a limit and you still exceed it, at that point, as the kids say, FAFO.
This is exactly the type of scenario I was thinking about.
> But this error is silent.
You’re conflating two different systems; the API and the client that accesses it. These are different systems worked on by different people in different organisations.
This error is not silent at all. It’s correctly flagging the error as soon as it happens. The client is told in no uncertain terms that the call failed. This is good engineering. Hiding a problem only makes it harder to detect, which prolongs time to fix. Errors should be flagged immediately and loudly. For an HTTP-based API that means responding with a 4xx class error and a response body containing an informative message. So – not silent at all.
Separately, in a different organisation, some client developer has screwed up by ignoring the response from an API call. In the general case, you can’t fix this externally. Bad developers who assume all their calls succeed and don’t check error conditions are going to make that mistake all over their code and their shortcomings are going to manifest as numerous bugs. Furthermore, whomever reviewed their code missed something obvious – these kinds of bugs are easy to spot in code review. So generally speaking, if you’re asking what I would do then the answer is nothing. I’m not responsible for fixing somebody else’s dysfunctional team in somebody else’s organisation and I’d just be scratching the surface by trying to work around just one of their bugs at the expense of my own service’s quality.
In this particular case, however, you can surface 4xx class errors in a number of ways if you really wanted to, e.g. email the client a daily report of all client errors. But you should be designing your API so that correct usage provides a robust solution and that means rejecting invalid data immediately.
If somebody is sending a disclosure in a separate notification to the message that needs the disclosure, they are already running afoul of the law because they have no way of ensuring the disclosure is delivered.
The only way for an API like this to operate correctly is to reject invalid messages and not truncate. Silent data mangling is not acceptable.
This post is saying to avoid OAuth and use bearer tokens because OAuth is too complicated. I agree with OAuth being too complicated, but I don't really think bearer tokens are the solution either.
Now there are jwts, and passkeys, and all these other solutions, which from where I'm standing, just look like someone else's resume-builder that I'm going to have to understand in a few years in order to do something simple and unrelated.
We are already connecting over TLS, just have the client authenticate with a long lived asymmetric key. Let me see that key in whatever web framework; I should have access to it, same as an HTTP header. Then I'll stick it in the database, maybe hash it first, if it's big. It doesn't have to be harder than that, your identity is (the hash of) your public key.
The public key does need to be exchanged, along with a signature relating it to the current session. This is all handled by TLS, there is no need for the client to send the key in the application data.
> You still need to generate it and distribute it to the user
This approach avoids distributing secret key material at all. Private keys should ideally never move. They are generated randomly, used to derive the corresponding public key, and then persisted as appropriate. The public key is sent around to other parties.
If you’re suggesting to just store a cert thumbprint that means a db call on every request - no different than just a secret token.
I wish implementing auth wasn't so complicated now, but the idea that we can just decide to make it so by doing it another way is a fantasy. That ship has sailed.
From personal experience and those of family members, I'd say we maintain many accounts leftover from bespoke sign-in processes and one-off services. That hasn't turned us from signing-up for these services, in spite of cumbersome onboarding (it's just one of those unfortunate 'hidden costs' of the web).
I say this, because we recently went through the trouble of deactivating many of those accounts that were no longer in use. Although the ship may have sailed for alternatives, it will be some time before legacy sign-ons will have been phased out, if at all.
Drive is still useful sometimes, mostly because my wife is addicted to it, but for a lot of the stuff I was using this for, I now use my Remarkable (and their SaaS).
Photos is one of those where I debate the usefulness of any of it. 99% of my photos are shite and not worth saving. The 1% that are, I need to curate a lot more than I do, and the curation is the problem not the storage. Ideally I'd like 2 buttons on the camera, one for "take this photo and delete it in 24 hours" and the other "take this photo and add it to my long-term storage with $this tag"
Getting third parties off Google login and on to email/password (using a paid password manager) was pretty simple, though I still find occasional sites that are just "welcome back <gmail address>!" and I have to work out how to get that changed. Haven't had any huge problems with that so far, apart from the usual "change email address, try to login again, get wrong password, reset password, 745395739 captcha attempts, generate password in password manager, save new password, login successful finally" dance.
You have an API provider service, an API client service, an a human telling the provider that the client should be allowed to access a limited subset of data on the provider service.
That's fine in cases where you actually need three parties involved, e.g. an API that allows access to a user's photo library stored in the cloud... but that's not every situation and if you can avoid it, you should.
A simple two party system, where the API provider authorises the API client, is vastly simpler. And just as secure (more secure even, because complexity is the enemy of security).
> Assume a human will read them
Yes, but... on the flip-side, make sure there's also some sort of code that the machines can read. Depending on the framework you started with, one or the other may be the more-obvious one to default to. But if you don't have both, you either have humans looking at cryptic codes (the point the article was making) or you have machines parsing through human-readable error messages with the risk of breaking backward-compatibility if you try to fix a typo in an error message or later decide to reword it or include more pointers.
This makes a rather specific assumption that there is no L7 load balancer or the like in front.
https://docs.aws.amazon.com/STS/latest/APIReference/API_GetA...