That seems hilariously aggressive to me, but her server her rules I guess.
That seems hilariously aggressive to me, but her server her rules I guess.
NixOS defaults to refresh frequency of every 5 minutes[0] (0_0).
I had noticed some blogs blackholing me before, but never quite made the connection.
So now it is configured to fetch every 12 hours. I believe that is fair.
[0] https://github.com/NixOS/nixpkgs/blob/d70bd19e0a38ad4790d391...
Why did you view the XML file directly?
There are many reasons why I do personally this.
1- Check that the link actually loads and works!
2- See how much content and does it contain last n or all feed history by default
3- To see if the feed gives summary or full content of posts
4- Just for curiosity like in this case I wanted to see what is this feed that prompted a blog post that reached HN front page.
It is usually a superposition state of those reasons. But this is why it is aggressive limit and I know it her server her rulee but this wasn't pleasant experience for me as an end user. I was just sharing my experience.
That's certainly a bit more effort to implement, though, and the author night not think it's worth the time.
The correct thing to do here is put a caching layer in front so that every feed reader isn't simultaneously hitting the origin for the same content. IP banning is the wrong approach. (Even if it's only a temporary block, that's going to cause my reader to show an error and is entirely unnecessary.)
So it means 30 months of blog posts content in single request.
Sending 0.5MB in single rss request is more crime than those 2 hits in 20 minutes.
There are a lot of very valid use cases where defaulting to deny for an entire 24 hour cycle after a single request is incredible frustrating for your downstream users (shared IP at my university means I will never get a non-429 response... And God help me if I'm testing new RSS readers...)
It's her server, so do as you please, I guess. But it's a hilariously hostile response compared to just returning less data.
So provide a poor service to everyone, because some people doesn't know how to behave. That sees like an even worse response.
No standards need to be updated. The client software needs to be a better HTTP citizen.
Sounds like something that could be scored in the rss reader tests.
Clever.
Gosh darn, if only I could say "Hey, please only send me the data if it's been modified since I last requested it an hour ago" somehow.
That’s my kind of humor.
I might be getting old, but 500KB in a single response doesn't feel "light" to me.
500KB is horrible for RSS.
100 articles are not reasonable.
100 articles where most of them are 1+ year old is madness.
RSS is not an archive of the entire website.
The host is fine with sending 0.5 MiB once (the client should be aswell from both a bandwidth and storage point of view).
The host is not fine with sending 0.5 MiB every 20 minutes, which could be easily avoided if the client would use the mentioned "If-Modified-Since header".
Whole-article feeds end up become exactly that - a local archive of a blog.
72 requests per day is nothing and acting like it's mayhem is a bit silly. And for a lot of people would result in them getting possible news slower. Sure OP won't publish that often but their rate limiting is an edge case and should be treated as such. If they're blocked until the next day and nothing gets updated then the only person harmed is OP for being overly bothered by their HTTP logs.
Sure it's their server and they can do whatever they want. But all this does is hurts the people trying to reach their blog.
As I pointed out, her blog and rate limiting are an extreme edge case, it would be silly for anyone to put effort into changing their feed reader for a single small blog. It's bad product management.
72 requests per day per IP over how many IPs? When you start multiplying numbers together they can get big.
Correction - from my billing page it's $4.50 a month, from the resize page it is $6 so I'm guessing I am grandfathered in to some older pricing
Imagine the petabytes of data transferred through the internet saved if a couple RSS clients added that method.
If OP reduced 30 months of posts in rss to 12 months then this 13mb would be 5mb a day.
Using Cloudflare free plan and this static content is cached without any problem.
I think making clients behave correctly is much more sustainable solution, although we could do better than doing so at the cost of the end users.
And since I've ran integrations that connected over 500 companies. I know what a rouge client actually looks like and 72 requests per day and I wouldn't even notice.
It is free and easy to scale this kind of text based blog.
Hey, I think you mistook HN for Reddit.
A similar problem arise from the increase in AI scraper activities. Talking to other SREs the problem seems pretty wide spread. AI companies will just hoover up data, but revisit so frequently and aggressively that it's starting to affect the transit feeds for popular websites. Frequently user-agents wouldn't be set to something unique, or deliberately hidden, and traffic originates from AWS, making it hard to target individual bad actors. Fair enough that you're scraping websites, that's part of the game when your online, but when your industry starts to affect transit feeds, then we need to talk compensation.
Like how "unlimited traffic, but will slow down to 1bps if you use more than 100gb in a month" is technically "unlimited traffic".
But for all intents and purposes, it's limited. And 429 are blocking. They include a hint towards the reason why you are blocked and when the block might expire (retry-after doesn't promise that you'll be successful if you wait), but besides that, what's the different compared to 403?
But I’m not, because they’re not blocking me. They’re asking my client to slow down. Neither AWS nor Rachel’s blog owes me unlimited requests per unit time, and neither have “blocked” me when I violate they policies.
See when you're trying to be pedantic and all about semantics, you should make sure you've crossed your Ts and dotted your Is.
> Block – AWS WAF blocks the request and applies any custom blocking behavior that you've defined.
from https://docs.aws.amazon.com/waf/latest/developerguide/waf-ru...
And my favourite
> Rate limiting blocks users, bots, or applications that are over-using or abusing a web property. Rate limiting can stop certain kinds of bot attacks.
From CloudFlare's explainer https://www.cloudflare.com/learning/bots/what-is-rate-limiti...
Every documentation on rate limit will include the word block. Because that's what you do, you allow access for a specific amount of requests and then block those that go over.