Scaling your API with rate limiters
stripe.com
stripe.com
There is also a corresponding set of code examples at the bottom that might be of interest to you: https://gist.github.com/ptarjan/e38f45f2dfe601419ca3af937fff...
But the simplification is that you do not need to write any application code. I can very easily have certain routes or domains limited differently by updating the nginx config. Especially in a microservices world I am trying to avoid having to update N services to get all of them to be rate limited.
The whole thing is 27 unique files, 1068 lines of Go, according to cloc.
[1]: https://github.com/xuqingfeng/caddy-rate-limit/blob/master/c...
Rate limiting is not meant to be a hard defence against denial of service attacks - that is an entirely separate engineering problem, where even a reverse proxy is not enough protection when you're facing a network-level flood. The primary purpose of rate limiting isn't to prevent hitting the application at all - it's to prevent hitting the heavy logic that starts when the controller is invoked. As long as your application's bootstrapping layer is lightweight, there is no problem with leaving rate limiting as part of the application.
For example, limiting traffic by IP address via reverse proxy is simple, but it seems like it would be more difficult to limit by request priority.
I was surprised when the article revealed it was middleware, and suddenly that made a lot more sense and seemed easier because it no-longer requires duplicating application logic to understand the content of requests. Middleware definitely seems like the better approach to me after these considerations.
What kinds of methods would one use to solve the problem of needing to parse and understand incoming requests using the reverse proxy method?
From what I saw, HAProxy appears to be even more powerful. Its ACL concepts are completely able to create rules based on headers, IPs in the request, etc. and you can compose them into larger ACLs.
With the example of request priority, if you can determine the priority by it's URL or a header, let's say, you can achieve this with nginx. But if you need to look the user up in a DB and see how much they're paying you, you obviously have to do it in the application.
Does this mean every rate limited request is hitting redis?
I had this problem years ago, and I used varnish to offload the rate limited traffic, which scaled very well then. Here is a blog post I wrote about it: http://blog.dansingerman.com/post/4604532761/how-to-block-ra...
(it was written directly in response to a 'One of your users has a misbehaving script which is accidentally sending you a lot of requests' issue)
So yes, our Redis traffic is strictly higher than our whole API requests per second. This is actually pretty easy since Redis scales pretty well horizontally (the keys shard well) in addition to being able to handle many orders of magnitude more traffic than a normal web stack per machine.
Thanks for linking to your blog post. I like the simplicity of not having to enter the web stack at all. I'm not very familiar with varnish, how do you end up doing the actual rate limiting algorithm? I see you making a hash but not dripping tokens into a bucket.
*This was 2011. The 429 response code wasn't defined until 2012.
ex: https://github.com/blendlabs/go-util/blob/master/collections...
If you look at the concurrent request limiter I do indeed keep all the timestamps there in a Redis set. That one was more error prone to write in practice as I often would accidentally hit Redis storage limits.
Check out my example: https://gist.github.com/ptarjan/e38f45f2dfe601419ca3af937fff... on line 23 with
local filled_tokens = math.min(capacity, last_tokens+(delta*rate))
that adds in any tokens that should drip into the bucket since the last check.yeah the algo is pretty standard, what we found is some edge cases get super weird, namely if you check on regular intervals you'll get false positives, vs. bursty calls
Similarly Im not sure about your "atomicity requirements." Is percieved accuracy driving the centralized store and single token semantics? Time skew alone is going to drive 0.1-1% innaccuracy in your rate limiters unless youre in crazy town doing 1pps or ptp everywhere.
Some food for thought is consolidating "rate" and "concurrent" limitations. Youll frequently see the same dichotomy in pps and bps based limits. If you can derive a consistent unit of work you can use a single bucket. Ie 10 requests that take 1 cpu second each have the same "cost" as 1 request at 10 cpu seconds. In the macro queue theory has them as the same impact to the system.
You "fleet usage" case, and a few other comments, reads like a simulacrum of weighted queues. Expanding it to weighted priority queues is a really powerful abstraction. Imagine requests which are over their base rate limit are not discarded, but actually marked as the lowest priority. That allows you to do generalized strategies like rejecting requests (load shedding) based on queue depth (latency) and FIFO/FILO.
PS: everyone reading _please_ respond with a 400 series if the request fails due to client request properties like rate. It really really really sucks to try and guess if 503 means "I" was rate limited and should back off, or "you" failed and the request should be immediately retried.
PPS A tactic to consider is returning the estimated recharge time when a request is throttled. A good client can use that to intelligently back of and achieve homeostasis with your server resources.
One would think that this would be a standard feature in web servers.
What were the hard kinks you had to figure out when coding this system?
I'm a fellow developer, would be interested in hearing you describe your system in more detail.
Did you do anything to mitigate the scenario where multiple users are behind the same IP address? With this approach I would worry about locking out all users in a NAT when a single user misbehaves.
It's also scary to me to process something "last" - under high enough load you might never get back to the (potentially briefly) abusive user? Did you attempt to guarantee some minimum passthrough rate even for misbehaving users?
This allow us to define the parameters of the token bucket on the service instead of the application like:
'requests':
per_second: 100
size: 200
Then in the app is like conformant = limitd.take('request', ip, 1)
I found the token bucket alg to be useful in a lot of scenarios not only for rate limiting but also for any kind of event debouncing. A common example: lets say you want to email a user everytime they trigger some condition on the system but you dont want to send the same mail more than once a day.Our project is opensource, we still working on it:
I just wonder, do you use rate limiters just for external API or also for internal API of your microservices?
One difference for internal rate limits is we set alerts whenever they fire so that we can track down the client and see what they were doing. It is much easier to do that when you own both sides of the pipe.
So the answer is probably a yes?
Also, do not forget about Quotas which usually comes along with Rate limits. Modern API gateways can handle so many stuff for you and help with API scaling.
Primarily I use this for idempotency and throttling things like usage events. But you can also use it for locking and concurrency control.
I ended up getting maybe... 3 or 4 well-behaved visitors / day for the first two years.
I thought it appropriate to plug my java based rate limiting library that implements the token bucket algorithm as mentioned in the article
Wondering how the two words "scale" and "limit" go together. My only experience right now with someone's API and rate limiting is Cloudinary. They give you 500 requests/hr (free). Which can be a lot or nothing at all.
No sorry it does make sense, don't allow one person to exhaust/consume your resources.
What about using Go? I hear crazy stuff like going from 2000 servers to 2, and doing 25,000 requests per second. Or is this a bandwidth concern?
For the case here, where the client insisted on keeping the 500/hr limit, I cached the query-results with a database as the values don't really change. But if they did change in the future that could be a problem, so you'd have to update the "cache".
Regarding handling overflow, that's something I haven't done yet myself, still stuck in the LAMP days, and have not done something like bench marking my server to see how many concurrent connections/requests it could handle.
note: LAMP isn't an excuse, I'm just making a note that I'm behind in using other technology that could be better. But I've heard of modular Apache configs/routing... stuff beyond me at this point.
Thanks in advance.
For my example in https://gist.github.com/ptarjan/e38f45f2dfe601419ca3af937fff... you would just set REPLENISH_RATE to be a different value for different users.
I'm guessing here but I'm guessing Stripe would actually only enforce these limits when they have some infrastructure issue or emergency and in most cases would allow all non-abusive uses of their API through.
Twitter and other such services could use much lower rate limits because its OK if a user is unable to post a new tweet for a few seconds.
If you are building a developer product where your API is your business, using rate-limiters is akin to preventing customers from giving you money. If your product is an API, you should encourage usage. The more usage you get, the more success both you and your customer will have. It's a win-win.
Because of this, I believe implementing rate-limiting strategies result in not only poor-DX for the product, but also a loss of trust (if these limits are in place to prevent your backend from failing, what other things do I need to worry about while using your product?), AND most importantly, they result in loss of business for both the API company and the customer.
IMO, if you're an API company, and you can't handle bursts of traffic from your customers, you should work on improving your backend and stop wasting time messing around with implementing patterns like this. It's a lose-lose situation for you and your customers.
I was really inspired by this recently after spending some time @ Twilio. Jeff (and the rest of the Twilio team) are hardcore about their API first product / thinking. They have a motto which is something along the lines of this: if you and your customer are both more successful and both make more money when the API usage goes up, do whatever you can to get out of the way and let the API be used as much as possible. I thought that was an awesome approach to take.
The way most smart companies do it is totally sufficient. Low but reasonable entry-level rate limit with a simple support contact to get it lifted. Maybe you're not aware but that's the exact model that Twilio uses[0].
[0] - https://support.twilio.com/hc/en-us/articles/223183648-Sendi...
I'm not against allowing the customer to set rate limits, etc. What annoys me in general are services where I want to use them and sometimes do burst traffic for a variety of reasons: but can't for an artificial reason :(
And, you don't make more money because someone's buggy script is hammering you with 1000 GET requests a second (for example).
I'm surprised this isn't more obvious, especially to a company like Stripe that is built on one of the slowest programming languages/platforms in existence. It isn't hard to come up with a list of languages that would easily provide an order of magnitude more capability per server (yes, maybe compute is not their bottleneck, but I've never seen an RoR setup that was bottlenecked by anything other than RoR). And that not only makes it cheaper for the baseline, but it also makes it cheaper to overprovision for safely handling spikes, and both cheaper and faster to scale when needed.
And like you said, not responding to demand is just preventing your customers from giving you money...but it is worse. If you are a growing company and one of your badly needed services is your bottleneck, that service is gonna be the first one to go. API limiting should be considered an attrition risk.
It's an European company that's 10 times the size and has lower fees.
Stripe having a harder time on that side of the Atlantic ;)
Scaling an API of any real value is NOT trivial, and struggling to scale an API to meet user demand does NOT necessarily mean that the backend was poorly designed. This is a naive generalization that is hazardous to the industry. Please don't spread it.
Here are some reasons why a lack of rate limiting / user auth is practically negligence. There are more, to be sure. I have experience operating a customer facing API for Bazaarvoice, so I think I know what I'm talking about. (We do thousands of requests per second and power reviews for the likes of Walmart, Best Buy, and 4,000 other retailers and brands worldwide.)
* Multi-tenancy * * client A over extends and causes client B to be unable to use the API * * client A needs scale independent of clients B-F * monitoring * * suddenly a client is making fewer than usual calls, why? * * suddenly a client is making more than usual calls, why? * billing * * want more requests / second? Upgrade your contract * * it's easier to measure how much I should charge customers per request or type of request, when I can see the rates of those requests and what it costs me * security * * DDOS attack? Start by setting the limit to nothing, or rejecting the requests * * leaking API auth info is less dangerous, if it happens
I think some other sibling comments mentioned other great reasons. The takeaway is that a valuable API will most likely be difficult, expensive or both difficult and expensive to scale, and rate limiting is extremely important.
Using a leaky bucket algorithm, and a per-customer bucket, I think it's possible to build "fair" systems that also improves the performance.
That is, you can run the system with a higher total transactions per seconds, just by queuing "simultaneous" requests a few milliseconds, as they will complete quicker.
The reason is probably that it's reducing contention and levels out the resource usage.
I thought it would be a feature in almost all web servers, since it's been known "since forever" in the telecom world, but I have not seen it. (Have not looked specifically either, so maybe there are good support for this everywhere and I missed it...)
Load shedding helps mitigate failures of insufficient capacity. When you need to prioritize some things over others load shedding helps here.
Counter-argument: That's not a reason not have these things, that's why you should think about implementing patterns like this before a unforeseen situation happens where they are useful. So they can limit the pain while you fix whatever causes the current load problem.
I agree when it comes to strict limits like "a single customer can only do X requests per hour", but that's not all rate limiting is good for.
Improving performance can be significantly more expensive and time-consuming than implementing simple rate limits. A realistic scenario is something along the lines of:
- 99% of time your request rate is N requests/min or lower
- 1% of time your request rate exceeds N requests/min, which could cause service degradation
You can deploy infrastructure to handle the 99% case, slap rate limits in front of the service, and sleep well at night. It's often not worth it to pay for additional infrastructure, spend time optimizing, etc for that 1% case.
As an API user, the way to think about rate limits is as a form of protection from other (misbehaving) users. Everyone is going to be upset when one user's script has a tight-loop firing at 100000 requests/sec and then the API becomes lethargic, throws errors, or goes down all together.
If you are counting requests/minute, you're seriously averaging your peak load and you're about to epicly fail in production.
For example, I inherited a relatively poorly performing service with a per-user limit set to 60 req/min. As the service operator, I have the choice of setting either a rate limit of 60 req/min or 1 req/sec. The former (60/min) leaves you open to per-user spikes of up to 60 req/sec. This is a real risk: 10 users together _could_ produce a spike of 600 req/sec.
Still, we went with the per-minute rate limit. Why? Those multi-user bursts are fairly improbable (most of our users were well-behaving) and queueing let us manage these isolated spikes gracefully. The per-minute rate limit is a bit more flexible from a user perspective (reward the good users) and it still combats the problem of sustained heavy load on the service, which is the true danger (stop/limit the "bad" users).
Performances must never be accounted in request/min. It's only requests/s that matter, because your performances are defined by the peak load you can take (which is many times the minute average).
Limits should be on more than a few seconds to support short peaks, yet throttle quickly if they persist.