Adaptive rate limiting is a game-changer
blog.fluxninja.com
blog.fluxninja.com
Just click around a bit:
- https://github.com/fluxninja/aperture
- https://docs.fluxninja.com/use-cases/adaptive-service-protec...
Note: I am one of the authors' of this project.
The "premature optimization is the root of all evil" misquote really caused a lot of damage to favor growth and profit making instead of better software. It seems programmers have a general habit of just implementing things in the most naive and cheap way and just solve any problem with expensive hardware.
We really need a new trend of fast and minimal software because climate change and a finite planet will not allow us to follow Wirth's law for very long.
Optimization should have measurable improvement and the need for it should be demonstrable. Neither of these is harder than writing a benchmark and/or performing some simple math on data measured for that purpose before going about the optimization.
> "The real problem is that programmers have spent far too much time worrying about efficiency in the wrong places and at the wrong times; premature optimization is the root of all evil (or at least most of it) in programming."
and another instance of the same:
> "Programmers waste enormous amounts of time thinking about, or worrying about, the speed of noncritical parts of their programs, and these attempts at efficiency actually have a strong negative impact when debugging and maintenance are considered. We should forget about small efficiencies, say about 97% of the time: premature optimization is the root of all evil. Yet we should not pass up our opportunities in that critical 3%."
https://en.wikiquote.org/wiki/Donald_Knuth#Computer_Programm...
I'm not really seeing much of a "misquote" to be honest. I have always understood this quote to mean that spending too much time micro-optimizing things unnecessarily will waste your time and become a burden to you later. Instead, optimizations should only be performed when the need for optimization becomes manifest.
for i in list_of_files; do
run_intensive_process &
sleep $(awk '{print $3}' < /proc/loadavg)
done
As more processes start upping my load average, it sleeps longer before running each one in the background. I guess GNU parallel can do this sort of thing with a load-sensitive command line arg, but I find this pattern easy to remember, and parallel's decidedly not.Works for me.
https://linux.die.net/man/1/at
I don't think at is very heavily used on modern systems, so it's quite possible that it doesn't actually have very useful behaviour here.
at(1) runs jobs after a specified time. (A common mistake is to presume that "at" means "will run when specified" when in fact it means "will run no earlier than specified", and depending on system characteristics, could run substantially later.)
batch runs jobs ... as system load permits, without any time specification, in a FIFO queue.
There's also cron, the third Unix scheduler, which executes a set of specified tasks on a regular schedule. All three are related.
As far as I know, cron doesn't pay attention to load average, but I would love to be corrected.
All of this feels like a very crufty and highly neglected part of Unix, which is a shame. There's an alternate history of Unix where it's all about batch jobs and datagram Unix domain sockets.
It does not. Up to whatever you run from cron to deal with that.
On many systems they'll share an integrated manpage.
cat list_of_file | parallel --load 100% run_intensive_processThey would just have to limit requests by IP and thats it. Say 10 tweets per IP per day. For requests which do not come from logged in users.
I can't see how scrapers would be able to come up with enough IPs to make a dent in Twitters costs.
So apparently not enough false positives for them to care.
Secondly, if they would, they could simply rate limit by the first x chars of the address.
And for the "simply log-in": It is even more expensive (way more expensive) to get good email adresses as it is to get good IPs. Email providers have very tough checks these days before they give you one.
I deal with these constantly for our clients. It’s amazing how smart they function to get around limits (thousands of IPs and valid recent user agents) but how they miss easy things like using the same cookie to submit 1000s of different entries with obvious purchased lists of names/emails, many of whom are deceased.
Reality would be that limit would have to be much higher.
https://twitter.com/SarahKSilverman
I check it ONCE per day, thats it. and for weeks every time it just prompts me to login. no thanks, done with Twitter.
The rest is the service provider's problem.
> They might prioritize serving the VIPs and slow down the pace when there is an influx of newcomers.
Or they might prioritize people who are about to convert and deprioritize them once they pay.
Even in networking, the classic application domain of adaptive rate limits (like TCP's algorithms), core networks are built to offer predictable performance. Edge networks, like public WiFi, that offers best-effort performance with adaptive rate limits, are often met with frustration and even anger the moment those limits kick in.
Google amp used to have a ridiculous 4 second timeout waiting for the font