Feed readers which don't take "no" for an answer
rachelbythebay.com
rachelbythebay.com
How would they know? Well, because Google, in its omniscience, started to downrank them for faking views with bots (which they do not do): it shows bot percentage in traffic stats, and it skyrocketed relative to non-bot traffic (which is now less than 50%) as they started to fall from the front page (feeding the vicious circle). Presumably, Google does not know or care it is a bot when it serves ads, but correlates it later with the metrics it has from other sites that use GA or ads.
Or, perhaps, Google spots the same anomalies that my friend (an old school sysadmin who pays attention to logs) did, such as the increase of traffic along with never seen before popularity among iPhone users (who are so tech savvy that they apparently do not require CSS), or users from Dallas who famously love their QQBrowser. I’m not going to list all telltale signs as the crowd here is too hype on LLMs (which is our going theory so far, it is very timely), but my friend hopes Google learns them quickly.
These newcomers usually fake UA, use inconspicuous Western IPs (requests from Baidu/Tencent data center ranges do sign themselves as bots in UA), ignore robots.txt and load many pages very quickly.
I would assume bot traffic increase would apply to feeds, since they are of as much use for LLM training purposes.
My friend does not actually engage in stringent filtering like Rachel does, but I wonder how soon it becomes actually infeasible to operate a website with actual original content (which my friend co-writes) without either that or resorting to Cloudflare or the like for protection because of the domination of these creepy-crawlies.
Edit: Google already downranked them, not threatened to downrank. Also, traffic rose but did not skyrocket, but relative amount of bot traffic skyrocketed. (Presumably without downranking the traffic would actually skyrocket.)
Some of the new traffic did come directly from Tencent data center IP ranges and reportedly those bots signed themselves in UA. I can’t say whether they respect robots.txt because I am told their ranges were banned along with robots.txt tightening. However, US IP bots that remain unblocked and fake UA naturally ignore robot rules.
Well, of course not, since the service is illegal.
New administration is going to be monopoly friendly.
I was honestly pleased that Gaetz was nominated for AG solely because he's big on antitrust. Or has been.
https://www.bbc.com/news/world-us-canada-57754435
Note that he's still talking about breaking up tech companies but not... X? (Surely that will resume once he and Elon have a falling out)
EDIT: Direct quote - "The internet's hall monitors out in Silicon Valley, they think they can suppress us, discourage us. Maybe if you're just a little less patriotic. Maybe if you just conform to their way of thinking a little more, then you'll be allowed to participate in the digital world,"
This isn't an attempt to ensure freedom from monopoly, this is an attempt to enforce partisan control of the message, weaponizing the idea of free speech using force.
I can assert that the 'common public square' idea central to freedom of speech is disappearing, and that this is a bad thing, but that's not what this man has been arguing or why this man has chosen this issue.
You're thinking of it wrong, the seeds of the thinking error are here: "I wonder how soon it becomes actually infeasible to operate a website with actual original content".
Bots want original content, no? So what's the problem with giving it to them? But that's the issue, isn't it? Clearly, contextually, what you should be saying is "I wonder how soon it becomes actually infeasible to operate a website for actual organic users" or something like that. But phrased that way, I'm not sure a CDN helps (I'm not sure they don't suffer false positives which interfere with organic traffic when they intermediate, more security theater because hangings and executions look good, look at the numbers of enemy dead).
Take measures that any damn fool (or at least your desired audience) can recognize.
Reading for comprehension, I think Rachel understands this.
Many are trying to index the web for whatever reason. By feeding them a Library of Babel, you can clog up their storage with noise.
The important part here is to do this chaotically. The worst sites to scrape are buggy ones. You are, in essence, deliberately following bad practices in a way real users wouldn't notice but would still influence bots.
Lmao!
104.28.42.8 - - [21/Dec/2024:13:58:35 -0800] consulting.m3047.net "GET /apple-touch-icon-precomposed.png HTTP/1.1" 404 980 "-" "NetworkingExtension/8620.1.16.10.11 Network/4277.60.255 iOS/18.2"
104.28.42.8 - - [21/Dec/2024:13:58:35 -0800] consulting.m3047.net "GET /favicon.ico HTTP/1.1" 200 302 "-" "NetworkingExtension/8620.1.16.10.11 Network/4277.60.255 iOS/18.2"
104.28.42.8 - - [21/Dec/2024:13:58:35 -0800] consulting.m3047.net "GET /dubai-letters/balkanized-internet.html HTTP/1.1" 200 16370 "-" "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_11_1) AppleWebKit/601.2.4 (KHTML, like Gecko) Version/9.0.1 Safari/601.2.4 facebookexternalhit/1.1 Facebot Twitterbot/1.0"
104.28.42.8 - - [21/Dec/2024:13:58:35 -0800] consulting.m3047.net "GET /apple-touch-icon.png HTTP/1.1" 404 980 "-" "NetworkingExtension/8620.1.16.10.11 Network/4277.60.255 iOS/18.2"
# dig -x 104.28.42.8
; <<>> DiG 9.12.3-P1 <<>> -x 104.28.42.8
;; global options: +cmd
;; Got answer:
;; ->>HEADER<<- opcode: QUERY, status: NXDOMAIN, id: 35228
;; flags: qr rd ra; QUERY: 1, ANSWER: 0, AUTHORITY: 1, ADDITIONAL: 1
;; OPT PSEUDOSECTION:
; EDNS: version: 0, flags:; udp: 1280
; COOKIE: 6b82e88bcaf538fc7ab9d44467685e82becd47ff4492b1be (good)
;; QUESTION SECTION:
;8.42.28.104.in-addr.arpa. IN PTR
;; AUTHORITY SECTION:
28.104.in-addr.arpa. 3600 IN SOA cruz.ns.cloudflare.com. dns.cloudflare.com. 2288625504 10000 2400 604800 3600
;; Query time: 212 msec
;; SERVER: 127.0.0.1#53(127.0.0.1)
;; WHEN: Sun Dec 22 10:46:26 PST 2024
;; MSG SIZE rcvd: 176
Further osint left as an exercise for the reader.I didn't realize this was an Apple thing, but that's fine. It changes the color of the horse and the name of the river, but the same road leads to the same destination.
1) There is a notion that Cloudflare is a content distribution network. The risk profile for a content distribution network is different from a VPN service. Now I know it's a VPN service (or is it?). Changes it from "seems weird and inappropriate" to "do I care about people relying on this? no, probably not". Cloudflare can't be arsed to provide reverse DNS for something which is clearly not part of their CDN, or is it?
1.5) Is it layer 2 or application? Cloudflare runs a CDN. Correct me if I'm wrong, but the CDN is a reverse proxy is it not? Is Cloudflare caching my website's content? Can they observe it? (It's surprisingly hard to find a solid explanation, but they talk about "proxies" and "decrypts the name of the website you requested" and none of that adds clarity, it makes it sound more like believe what we want you want to believe.)
2) I don't block incoming SYNs from Cloudflare (yet) the way I do with Amazon, and this traffic per se isn't going to trip any mitigations here. But not all of the traffic is as benign (and it's impressive that they're so technically savvy they don't need the CSS as noted elsewhere). Presumably those exit points are shared by multiple customers. Did I mention I block all incoming SYNs from Amazon?
With the logs you provided, they appear to be coming from within iMessage.
So when someone posts a link in iMessage it will fetch the favicon(s) and the html in order to generate a “preview” of the page with the title of the page and use one of the favicons. It doesn’t need to fetch any css files to do this.
Not saying bad actors don’t fetch css either, but the lack of it being fetched doesn’t mean that it’s a bad actor.
As for why CF don’t reverse DNS their IPs stating it’s iCloud private relay, well CF are not Apples only 3rd party egress provider (Akamai are also one that springs to mind). So if the number of providers can change at any time, the best source of information about valid egress providers is from Apple themselves.
But Apple do also publish these changes to geo-location databases for you to query, for example: https://www.ip2location.com/demo/104.28.42.8 lists it as iCloud Private Relay.
As for “are CloudFlare caching my site when ran through private relay?”, not 100% sure, I’ll have to check my own logs and cba’ed right now, but I don’t think so (it’s been a while since I ran tests on it to see how it behaved to be 100% sure right this minute.
But I think it would be silly of them if they did as they may not be aware of the what to cache and for who. Let’s say they cached /profile without knowing what the server is using to determine who the logged in user is, they may false cache-hit and leak data from a previous request. When they act as your sites CDN you explicitly tell them what to cache on, but when acting as a relay (either for apple or their own warp product) for a site they are not a CDN for they are missing this info, sure they could guess, but why risk being wrong?)
https://github.com/hroost/icloud-private-relay-iplist/blob/m...
(There is also a list of ranges on apples site, but I forget where…)
Edit: found it https://mask-api.icloud.com/egress-ip-ranges.csv
> 00:04:51 GET /w/atom.xml, unconditional.
> Fulfilled with 200, 502 KB.
> [...]
> A 20 minute retry rate with unconditional requests is wasteful. [...]
And If-Modified-Since makes a request conditional. https://developer.mozilla.org/en-US/docs/Web/HTTP/Conditiona...
> Advised (via Retry-After header) to come back in one day since they are unwilling or unable to do conditional requests.
But I think it's still unarguable that the post doesn't explicitly mention If-Modified-Since, which it's not obliged to do, but the mention of it here could be helpful to someone. So why fuss?
If you're interested in it, I highly recommend just reading from start to end. It's all quite interesting (including building out a whole test service for feed readers to get scored on their behavior)
That seems hilariously aggressive to me, but her server her rules I guess.
72 requests per day is nothing and acting like it's mayhem is a bit silly. And for a lot of people would result in them getting possible news slower. Sure OP won't publish that often but their rate limiting is an edge case and should be treated as such. If they're blocked until the next day and nothing gets updated then the only person harmed is OP for being overly bothered by their HTTP logs.
Sure it's their server and they can do whatever they want. But all this does is hurts the people trying to reach their blog.
As I pointed out, her blog and rate limiting are an extreme edge case, it would be silly for anyone to put effort into changing their feed reader for a single small blog. It's bad product management.
72 requests per day per IP over how many IPs? When you start multiplying numbers together they can get big.
And since I've ran integrations that connected over 500 companies. I know what a rouge client actually looks like and 72 requests per day and I wouldn't even notice.
If OP reduced 30 months of posts in rss to 12 months then this 13mb would be 5mb a day.
Using Cloudflare free plan and this static content is cached without any problem.
I think making clients behave correctly is much more sustainable solution, although we could do better than doing so at the cost of the end users.
Correction - from my billing page it's $4.50 a month, from the resize page it is $6 so I'm guessing I am grandfathered in to some older pricing
Imagine the petabytes of data transferred through the internet saved if a couple RSS clients added that method.
It is free and easy to scale this kind of text based blog.
Hey, I think you mistook HN for Reddit.
So it means 30 months of blog posts content in single request.
Sending 0.5MB in single rss request is more crime than those 2 hits in 20 minutes.
There are a lot of very valid use cases where defaulting to deny for an entire 24 hour cycle after a single request is incredible frustrating for your downstream users (shared IP at my university means I will never get a non-429 response... And God help me if I'm testing new RSS readers...)
It's her server, so do as you please, I guess. But it's a hilariously hostile response compared to just returning less data.
So provide a poor service to everyone, because some people doesn't know how to behave. That sees like an even worse response.
No standards need to be updated. The client software needs to be a better HTTP citizen.
Gosh darn, if only I could say "Hey, please only send me the data if it's been modified since I last requested it an hour ago" somehow.
Sounds like something that could be scored in the rss reader tests.
That’s my kind of humor.
Clever.
But I’m not, because they’re not blocking me. They’re asking my client to slow down. Neither AWS nor Rachel’s blog owes me unlimited requests per unit time, and neither have “blocked” me when I violate they policies.
See when you're trying to be pedantic and all about semantics, you should make sure you've crossed your Ts and dotted your Is.
> Block – AWS WAF blocks the request and applies any custom blocking behavior that you've defined.
from https://docs.aws.amazon.com/waf/latest/developerguide/waf-ru...
And my favourite
> Rate limiting blocks users, bots, or applications that are over-using or abusing a web property. Rate limiting can stop certain kinds of bot attacks.
From CloudFlare's explainer https://www.cloudflare.com/learning/bots/what-is-rate-limiti...
Every documentation on rate limit will include the word block. Because that's what you do, you allow access for a specific amount of requests and then block those that go over.
Like how "unlimited traffic, but will slow down to 1bps if you use more than 100gb in a month" is technically "unlimited traffic".
But for all intents and purposes, it's limited. And 429 are blocking. They include a hint towards the reason why you are blocked and when the block might expire (retry-after doesn't promise that you'll be successful if you wait), but besides that, what's the different compared to 403?
I might be getting old, but 500KB in a single response doesn't feel "light" to me.
500KB is horrible for RSS.
100 articles are not reasonable.
100 articles where most of them are 1+ year old is madness.
RSS is not an archive of the entire website.
Whole-article feeds end up become exactly that - a local archive of a blog.
The host is fine with sending 0.5 MiB once (the client should be aswell from both a bandwidth and storage point of view).
The host is not fine with sending 0.5 MiB every 20 minutes, which could be easily avoided if the client would use the mentioned "If-Modified-Since header".
A similar problem arise from the increase in AI scraper activities. Talking to other SREs the problem seems pretty wide spread. AI companies will just hoover up data, but revisit so frequently and aggressively that it's starting to affect the transit feeds for popular websites. Frequently user-agents wouldn't be set to something unique, or deliberately hidden, and traffic originates from AWS, making it hard to target individual bad actors. Fair enough that you're scraping websites, that's part of the game when your online, but when your industry starts to affect transit feeds, then we need to talk compensation.
The correct thing to do here is put a caching layer in front so that every feed reader isn't simultaneously hitting the origin for the same content. IP banning is the wrong approach. (Even if it's only a temporary block, that's going to cause my reader to show an error and is entirely unnecessary.)
Why did you view the XML file directly?
There are many reasons why I do personally this.
1- Check that the link actually loads and works!
2- See how much content and does it contain last n or all feed history by default
3- To see if the feed gives summary or full content of posts
4- Just for curiosity like in this case I wanted to see what is this feed that prompted a blog post that reached HN front page.
It is usually a superposition state of those reasons. But this is why it is aggressive limit and I know it her server her rulee but this wasn't pleasant experience for me as an end user. I was just sharing my experience.
That's certainly a bit more effort to implement, though, and the author night not think it's worth the time.
NixOS defaults to refresh frequency of every 5 minutes[0] (0_0).
I had noticed some blogs blackholing me before, but never quite made the connection.
So now it is configured to fetch every 12 hours. I believe that is fair.
[0] https://github.com/NixOS/nixpkgs/blob/d70bd19e0a38ad4790d391...
People often implement error handling using constructs like regexp matching on status codes, while with domain-specified errors it would be obvious what exactly is the range of possible errors.
Moreover, when people do implement domain errors, they just have to write more code to handle two nested levels of branching.
Well, good luck designing any standard app-independent protocol that works and doesn't do that.
And yes, you must handle two nested levels of branching. That's how it works.
The only improvement possible to make it clearer is having codes for API specific errors... what 400 and 500 aren't exactly. But then, that doesn't gain you much.
Perhaps put the app-specific part in the body of the reply. In the RFC they give a human specific reply to (presumably) be displayed in the browser:
HTTP/1.1 429 Too Many Requests
Content-Type: text/html
Retry-After: 3600
<html>
<head>
<title>Too Many Requests</title>
</head>
<body>
<h1>Too Many Requests</h1>
<p>I only allow 50 requests per hour to this Web site per
logged in user. Try again soon.</p>
</body>
</html>
* https://datatracker.ietf.org/doc/html/rfc6585#section-4* https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/429
But if the URL is specific to an API, you can document that you will/may give further debugging details (in text, JSON, XML, whatever).
Oh the horror. I would assume the practice is encourage by "RESTful" people?
Back-and-forth conversion is a very poor idea. It works for what was “internet resources” initially (basically files and folders), but later people stretched that on application data models and that creates constant issues because people naturally can’t understand the mapping, cause there’s none. This is not a good idea. Talk to http hosts with http and talk to your client with a language you designed specifically for talking to it. 200 vs non-200 is http level and orthogonal to in-service statuses.
The HTTP server, when it detects "URI handler not found" condition, builds an 404 HTTP response and sends it as a normal payload through the underlying connection instead of turning it into an TLS error packet or an RST packet on TCP level (that's the TCP's standard response for "port handler process not found", after all) or something silly like that, and that is absolutely fine, because the application-level (HTTP) error messages should be transmitted by the transport level (TLS/TCP) just as normal messages would.
The same reasoning holds just the same when we consider the usage of HTTP as a transport-level protocol for some higher-level RPC exchange. Yes, HTTP has some assortment of error codes that superficially look like they can be reused to serve as the upper-layer errors as well but that's a red herring.
There is a difference between /api/entity/123 and /api/search with a payload of 123, though.
The request couldn't be processed due to semantic errors... perhaps such as not being mapped to a handler :->
I suppose from the right point of view that could also be likened to a reverse proxy not being able to send the request on (502 bad gateway), but sane people would probably find that even more confusing.
There were also attempts to use 204 no content for "I successfully confirmed that what you asked for doesn't exist", but I think I managed to shoot those down.
But... why not throw a CDN in front of your site and focus your energy somewhere else? I guess every problem has to be solved by someone, but this just seems like a very strange hill to die on.
And she posts on it lots because she has a bunch of RSS clients pointed at her writing, because she's rather popular.
And she'd rather people writing this stuff just learn HTTP properly, at least out of professionalism, if not courtesy.
Hey, you might not, I might not, but we all choose our hills to die on.
My personal hill is "It's lollies and biscuits, not candy and cookies".
Yes it's been invented before, known as Feedburner, which was acquired & abandoned by Google.
Because this is how the open web dies - one website at a time. It's already near-dead on the client side - web browsers are not really "user" agents, but agents of oligopolist corporations, that have a stake in abusing you[1].
It's been attempted before with WAP[2], then AMP. But effectively, we're almost there.
This often requires to do lots of tests against the endpoint, which the server prohibits.
But are RSS reader devs willing to jump through such hoops?
I would claim that writing a (simple) RSS reader (using a programming language that provides suitable libraries) is something that would be rather easy for me, but setting up a caching layer would (because I have less knowledge about the latter topic) take a lot more research from my side concerning how to do it.
If I was doing local development on an RSS reader, I'd just download any atom.xml file that seemed relevant, then serve it locally using php -S or some other local HTTP file-system server.
That way I can hit that file a million times without bothering any remote server. Plus, then you've made your local development environment stable even in the face of internet outages.
And if you automate the local HTTP server setup a little bit, then you can even run local integration tests.
I dunno, it seems like if "write an RSS feed reader" is an easy problem, I'd expect "serve a file from my hard drive over HTTP on localhost" should also be an easy problem.
My RSS reader YOShInOn subscribes to 110 RSS feeds through Superfeedr which absolves me of the responsibility of being on the other side of Rachel's problem.
With RSS you are always polling too fast or too slow; if you are polling too slow you might even miss items.
When a blog gets posted Superfeedr hits an AWS lambda function that stores the entry in SQS so my RSS reader can update itself at its own pace. The only trouble is Superfeedr costs 10 cents a feed per month which is a good deal for an active feed such as comments from Hacker News or article from The Guardian but is not affordable for subscribing to 2000+ indy blogs which YOShInOn could handle just fine.
I might yet write my own RSS head end, but there is something to say for protocols like ActivityPub and AT Protocol.
Sure I could get DNS to point to my ADSL connection and set something up in my router so that my home computer can answer the webhook but then I can never turn my computer off. On top of that I have about one power outage a month.
It would also be non-proprietary to spend $5k on a server and $300 a month on colo costs but with AWS I can use what would be 1 cent of resources if I was optimizing that colo (with $10k of labor) and pay 10 cents for it (could really be spending up to $50 on a non-optimized colo, which is what I might have if I don't need to handle a billion webhooks a month) If I want to switch to Azure or some other service that could answer a webhook the labor involved is minuscule.
Surely a $15 per year vm would be enough to receive a few thousands POSTs per day (https://tinykvm.com/), or if that's not enough, paying less than $5 per month for a VPS is still cheap
Surely transposing a binary hosted on vendor A to vendor B is as minuscule in involved labor as configuring a webhook
I'm not trying to diss on your own reader which is an awesome thing to do, I'm just sad that resorting to a fully proprietary architecture is an automatism when considering a protocol that can't be more open than RSS
w/ AWS I get all kinds of charts or alarms free or very cheap. The lambda is about 20 lines of Python code, I was able to complete part of the project that I was uninterested in at the time very quickly. Later on it took maybe 2-3 hours to make a UI that would let me add, view and remove feeds from my reader's UI.
YOShInOn's recommendation engine and UI were a research project however that was high-risk and might not have worked so it wouldn't have made sense to develop a better head end.
A better head end is on the agenda today but that system has a lot of other problems such as a dangerously large database that needs to be pruned, O(N) algorithms that were fine for the first year, other problems on the tail end. And it competes with other systems.
My experience is I can set something like this up in AWS and just not think about it for years.
Hopefully you don't have some expensive code generating the feed on the fly, so processing overhead is negligible. But if it's not, cache the result and reset the cache every time you post.
Surely this is easier than spending the effort and emotional bandwidth to care about this issue?
I might be wrong here, but this feels more emotionally driven ("someone is wrong on the internet") than practical.
I never even considered the option or necessity. It's easy and cheap just to send everything.
I guess static generators with a apache style web server probably do, but I can't imagine any dynamic generators bother to try to save the small handful of bytes.
requests_per_day user_agent
283 Reeder/5050001 CFNetwork/1568.300.101 Darwin/24.2.0
274 CommaFeed/4.4.0 (https://github.com/Athou/commafeed)
127 Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36
52 NetNewsWire (RSS Reader; https://netnewswire.com/)
47 Tiny Tiny RSS/23.04-0578bf80 (https://tt-rss.org/)
47 Refeed Reader/v1 (+https://www.refeed.dev/)
46 Selfoss/2.18 (SimplePie/1.5.1; +https://selfoss.aditu.de)
41 Reeder/5040601 CFNetwork/1568.100.1.1.1 Darwin/24.0.0
39 Tiny Tiny RSS/23.04 (Unsupported) (https://tt-rss.org/)
34 FreshRSS/1.24.3 (Linux; https://freshrss.org)
Reeder is loading the feed every 5 minutes, and in the vast majority of cases it’s getting a 301 response because it tries to access the http version that redirects to https. At least it has state and it gets 304 Not Modified in the remaining cases.If I order by body bytes served rather than number of requests (and group by remote_addr again), these are the worst consumers:
body_megabytes_per_year user_agent
149.75943975 Refeed Reader/v1 (+https://www.refeed.dev/)
95.90771025 Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36
75.00080025 rss-parser
73.023702 Tiny Tiny RSS/24.09-0163884ef (Unsupported) (https://tt-rss.org/)
38.402385 Tiny Tiny RSS/24.11-42ebdb02 (https://tt-rss.org/)
37.984539 Selfoss/2.20-cf74581 (+https://selfoss.aditu.de)
30.3982965 NetNewsWire (RSS Reader; https://netnewswire.com/)
28.18013325 Tiny Tiny RSS/23.04-0578bf80 (https://tt-rss.org/)
26.330142 Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/84.0.4147.105 Safari/537.36
24.838461 Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/84.0.4147.105 Safari/537.36
The top consumer, Refeed, is responsible for about 2.25% of all egress of my webserver. (Counting only body bytes, not http overhead.)[1]: https://ruudvanasseldonk.com/writing [2]: https://github.com/ruuda/sqlog/blob/d129db35da9bbf95d8c2e97d...
I also design 2 new formats that no one (including myself) has ever implemented.
https://go-here.nl/ess-and-nno
enjoy
Disappointing combined with the various update sites it tries to contact every startup, which is completely unnecessary as well. Couple of times a week should be the maximum rate.
It's a Britticism (AFAICT) making inroads.
I'm making an observation. I try to spot language change. And "which" has been popping up in stupid places, at the expense of "that".
There is a real difference between "which" and "that", whether Joe Sixpack gives a damn or not. And nowadays a lot of people seem to be inappropriately overusing "which", and it seems to be driven by UK speakers.
Here's a sample explanation: https://www.grammarly.com/blog/grammar/which-vs-that
I should be making a collection of the most OTT examples.
That might be a sample of what you're thinking, but it isn't a correct explanation of anything. It's just some random mythmaking, which is par for the course from "grammar advice" websites.
Which and that are not distinguished in the manner that page wishes they were. That cannot be used (in the modern language) to introduce a nonrestrictive relative clause. Which can be used to introduce a restrictive or nonrestrictive relative clause.
> There is a real difference between "which" and "that", whether Joe Sixpack gives a damn or not.
There are real differences, but there are no differences as to the usage you flagged. The bigger difference is that which, being a pronoun, is part of a gender distinction (with who) between persons and nonpersons, whereas that, not being a pronoun, has no such distinction.
You might note, hopefully, that if feed readers were people, "feed readers who don't take 'no' for an answer" would be completely standard grammar. The same is also obviously true of "feed readers which don't take 'no' for an answer" when the feed readers aren't people.
CGEL [Cambridge Grammar of the English Language; both authors are from the UK] doesn't even bother to mention this myth, but the first example of a relative clause that it does give is He'll be glad to take the toys which you don't want. (Chapter 12 §2.1, example 1.i; page 1034)
If you want to play a grammarian on the internet, wouldn't it be better to know some grammar first?
> I try to spot language change.
> And nowadays a lot of people seem to be inappropriately overusing "which"
You're doing a remarkably terrible job; this change took place seven hundred years ago. Here's The Merriam-Webster Dictionary of English Usage:
> According to McKnight 1928 that was prevalent in early Middle English, which began to be used as a relative pronoun in the 14th century, and who and whom in the 15th.
> [...] By the early 17th century, which and that were being used pretty much interchangeably. Evans 1957 quotes this passage from the Authorized (King James) Version (1611) of the Bible:
>> Render therefore unto Caesar the things which are Caesar's; and unto God the things that are God's.
> During the later 17th century, Evans tells us, that fell into disuse, at least in literary English. It went into such an eclipse that its reappearance in the early 18th century was noticed and satirized by Joseph Addison in The Spectator (30 May 1711) in a piece entitled "Humble Petition of Who and Which against the upstart Jack Sprat That."
(entry for that [1]; page 894 of the 1993 printing.)
If anything, a 429 is a nice heads up. It could have been worse; she could have redirected those requests to a separate URL with an... unpleasant content, like a certain domain that redirects to I-don't-know-what whenever they detect the Referer header is from HN.
So my question to you is... ...why is it so anti-you, and why should it be different?
> Unconditional requests: at most once per 24 hour period.
> Conditional requests: at most once per 60 minute period.
(Source: calling `curl hxxps://rachelbythebay[.]com/w/atom.xml` twice)
I find this a particularly amusing comment, given that the main complaint is about feed reader's not sending conditional requests.
Learn2cache indeed, feed reader authors.
All feed readers/clients should cache responses when sending multiple requests the same day.