Thanks, Cloudflare.
Thanks, Cloudflare.
If you want to blame someone, blame the people who poisoned the well for everyone else.
How is this an "externality" of cloudflare's service? If anything, it is an "externality" of the server hosts.
But if we take a step back it's clear that Cloudflare is the entity here that can have the biggest impact. Spammers and customers are diffuse, Cloudflare isn't. How much of the blame they deserve, I don't really care, but theres a problem, they're in the best position to act, so they have a responsibility to. In my opinion, that's the best mindset for solving problems.
This is, to me, a bizarre notion. You're imposing some sort of moral edict upon cloudflare for providing an opt-in service to web hosts.
> they're in the best position to act
They have, and they decided that certain traffic requires additional consideration to distinguish from bad actors. Your refusal to cooperate is entirely on you.
Just because you don't like their solution doesn't mean that there is a better one. If you have better ideas, I'm sure they're all ears, as less intrusive is generally cheaper to implement.
Is it? Those who want to DDoS will always find a way, and meanwhile users with slightly odd hardware/software are being locked out. Admittedly the latter is a minority, but one of the key tenets of the Internet and what made it so successful was interoperability. This is, in some ways, even worse than (but somewhat of a following effect of) the effective browser monopoly.
Calling it "security" when it's really about "availability" is another deceptive misdirection, because the former is something that can more easily persuade the sheeple.
I don't really like to go to extremes, but fingerprinting clients and deciding access based on such should really be regarded as a moral equivalent to racial profiling.
To tie this back to the topic at hand, you're complaining that a service has decided your traffic resembles known patterns from bad actors, and is asking you to go through an extra step to access the content.
Are there better options? Maybe, but it's utterly asinine to compare what cloudflare is doing to racial profiling.
How do you figure? You don't think far more DDoS events would occur if it was easier and more effective?
Not really. We are actually pretty good these days at stopping DDOS attacks.
In other cases, people try more sophisticated attacks (e.g. posting random terms to a search page to avoid caching) and that’s more of a problem but it’s probably like 1% of the total traffic because it’s moved out of script kiddie territory into something where you need to have more skills and people don’t generally do that without a way to make money from it. One challenge with a DDoS in that regard is that it’s not subtle so your ability to wage an attack goes away relatively quickly without constant work replacing systems which are taken offline by a remote ISP.
This is like saying no lock will stop a thief. Cost and difficulty matter, and requiring a full browser increases the challenge for an attacker enough that some people will give up and others won’t be able to send as much traffic. That’s not perfect but speaking from experience a surprising fraction of people will give up after a naive attack fails.
Despite that I can visit websites where admins that set permissions not overly tough, I still run into enough blocks due to cloudflare that I'm considering investing time (ok for the last few years I've been really really lazy, I'm happy enough to copy and paste so I can manually write a nasty comment ) so that for each "bad load" the web site can be added to my disallow list. It might only save the smallest amount of wasted bandwidth but I guess it all adds up.
> Know your user is coming from an authentic device and signed application, verified by the device vendor directly.
If they manage to get wide adoption of something like this it seems like a very bad day for privacy, and a joyous day for advertisers, anyone who wants to prevent web scraping, Google and/or anyone who might want to make it hard to crawl the entire web…
I'm still having a hard time seeing how this doesn't eventually lead to a completely locked-down internet, where users can only use approved browsers and devices.
They have a list of steps for how a request would be made, where steps 2 and 3 are:
> 2. Safari supports PATs, so it will make an API call to Apple’s Attester, asking them to attest.
> 3. The Apple attester will check various device components, confirm they are valid, and then make an API call to the Cloudflare Issuer (since Cloudflare acting as an Origin chooses to use the Cloudflare Issuer).
In a theoretical future world where 99% of site operators have set this up for protection, and 99% of users are using approved browsers, how would one do something like...
Create a competitor to Google? You'd need to crawl the web for that. Would you imagine Apple or Cloudflare would gladly let your device request millions of tokens per hour? Or would that be throttled or disallowed entirely?
Use curl (or telnet, or [any other HTTP client]) to grab a page?
Use yt-dlp to download a YouTube video?
Scrape a bunch of data for an AI project? See this article from the front page where someone scraped a bunch of car listings from KBB and trained a model to estimate car prices and found some interesting results. https://blog.aqnichol.com/2022/12/31/large-scale-vehicle-cla... - would something like that be permitted under a system like this? Or might you need to own/rent an army of authorized devices with authorized browsers to do that experiment?
Here's the "you will be under our control" part of their scheme. Running any "unauthorised" software? Rooted/jailbroken? Certain "security" features disabled? Using third-party replacement parts? ... Social credit score too low? Too bad, you're now denied access.
For requests which are "reading" you just serve content, unless for some reason you only want human eyeballs to see your content.
In your proposed scenario, reading the publicly accessible contents of the web, there should be no problems. (Of course some percentage of sites will accidentally have required PAT at any time and be unscannable, but presumably they figure that out and fix it.)
Now for the good side: I, reluctantly, implemented a geolocation filter to control anonymous content additions to that service I was alluding to. I felt bad about it, but I also felt bad having to filter out content spam every day. It turned out that all my strange content spam came from one country, so I banned 143 million people from anonymous content creation for my convenience.
With PAT I can remove the national ban and let any "probably human" in.
I was going to cite the LinkedIn case where, last I had heard, the courts had decided that scraping was legal...
Headline [0] from April:
> Court rules that data scraping is legal in LinkedIn appeal
> LinkedIn has lost its latest attempt to block companies from scraping information from its public pages, including member pages.
... but upon googling it, I found a more recent [1] ruling :/
> LinkedIn prevails in 6-year lawsuit against data scraper
> The U.S. District Court for the Northern District of California sided with LinkedIn in its six year lawsuit against a firm that scraped data ...
So that sure puts a nail into the argument I was going to make. But still, while I think your use case lines up with the spirit of this kind of system, I think the reality is that it also would be used by every single site with a signup wall to kill off the archive.ph's of the world.
0: https://www.zdnet.com/article/court-rules-that-data-scraping...
1: https://iapp.org/news/a/linkedin-prevails-in-six-year-lawsui...
yeah, very few follow the original recommendation of showing captcha for ips that already showed signs of being bots etc.
but the vast majority shows captchas for absolutely anything. I can't even book a dmv appointment today without answering a gratutious one.
same will surely happen with PAT, especially because it is so easy for the implementer to shove it everywhere. people are lazy and dumb.
A: they won't. and that's the plan. not to mention that now those are the only two players (microsoft a late third) that can both attest you and profile you locally on the device for advertising profiling.
This is implemented in a privacy preserving way.
You will need to elaborate on that.
It's bad for freedom. Very, very very bad.
Talking about the "privacy" of what has been made publicly available makes no sense.
So are laws in the real world about stealing.
>Talking about the "privacy" of what has been made publicly available makes no sense.
Yes, it does. Users often wish to be able to delete or make something that was once public private. For example someone could post a picture of themselves on twitter. A year later they are no longer comfortable with having pictures of themself online so they go and delete them. Despite the user deleting them malicious scrapers will not delete them and keep those images. Another example would be setting your real name to your twitter name. Later you aren't comfortable using your real name so you change it away. Scrapers may still have your real name despite you wanting it to be a secret.
People also wish to be be able to do a lot of other things, but that doesn't make it right.
What becomes public history must remain immutable. Otherwise you're just going to encourage a state in which those who have the power to will destroy and rewrite the past to their advantage, to control the narrative over the population. The trendy phrase "right to forget" is effectively a "right to rewrite history".
It's interesting that you automatically call those wanting to preserve what could possibly be very important history "malicious scrapers".
I am going off of twitter's view. If you store tweets locally you must listen for when they get deleted and then delete them on your end too. If a scraper is breaking twitter's rules I consider that malicious scraper.
https://developer.twitter.com/en/developer-terms/policy#3.Up...
>If you store Twitter Content offline, you must keep it up to date with the current state of that content on Twitter.
https://www.ietf.org/archive/id/draft-ietf-rats-tpm-based-ne...
For example, this alleged "head-of-line blocking problem" that HTTP/2 purportedly "solves" was never a problem of HTTP outside of a specific program, the graphical web browser, the type of client that tries to pull resources from different domains for a single website. Not all programs that use HTTP need to do that.
For instance I have been using HTTP/1.1 pipelining outside the browser for fast, reliable information retrieval for close to 20 years. It has always been supported by HTTP servers and it works great with the simple clients I use. I still rely on HTTP/1.1 pipelining today, on a daily basis. Never had a problem.
There are uses for pipelining besides the ones envisioned by "tech" companies, web developers and their advertiser customers.
The maintainer of a popular webserver has suggested HTTP/2 is slower than HTTP/1.1 for file download.
https://stackoverflow.com/questions/44019565/http-2-file-dow...
As I stated, I use HTTP/1.1 pipelining every day. I use it for a variety of information retrieval tasks, even retrieving bulk DNS data. To give an arbitrary example, sometimes I will download a website's sitemaps. This usually involves downloading a cascade of XML files. For example, there might be a main XML file called "index.xml". This file then lists hundreds more sitemap XML files, e.g., archive-2002-1.xml, archive-2002-2.xml, containing every content URL on the website beginning with some prior year all the way up to the present day. Using a real world example, index.xml contains 246 URLs. Using HTTP/1.1 pipelining I can retrieve all of them into a single file using a single TCP connection. Then I retrieve batches of the URLs contained in that file, again over a single TCP connection. Many websites allow thousands of HTTP requests HTTP/1.1-pipelined over a TCP single connection, but I usually keep the batch size at around 500-1000 max. Of course I want the responses in the same order as the requests.
The process looks something like this
ftp -4o 1 https://[domainname]/sitemaps/index.xml
yy030 < 1|(ka;nc0) > 2
yy030 < 2|wc -l
1337855
1337855 is the number of URLs for [domainname]. Content URLs, not Javascript, CSS or other garbage.yy030 is a C program that filters URLs from standard input
ka is a shell alias that sets an environment variable that is read by the yy025 program to indicate an HTTP header, in this case the "Connection:" header set to "keep-alive" not "close" (ka- sets it back to close)
nc0 is a one line shell script
yy025|nc -vv h1b 80|yy045
yy025 is a C program that accepts URLs, e.g., dozens to hundreds to thousands of URLs, on stdin and outputs customised HTTPh1b is a HOSTS file entry containg the address of a localhost-bound forward TLS proxy
yy045 is a C program that removes chunked transfer encoding from standard input
To verify the download, I can look at the HTTP headers in file "2". I can also look at the log from the TLS proxy. I have it set configured to log all HTTP requests and responses.
Is this a job for HTTP/2. It does not seem like it.
This type of pipelining using only a single TCP connection is not possible using curl or libcurl. Nor is it possible using nghttp. Look around the web and one will see people opening up dozens, maybe hundreds of TCP connections and running jobs in parallel, trying to improve speed, and often getting banned. As with the comment from the Jetty maintainer, I suspect using HTTP/2 would actually be slower for this type of transfer. It is overkill.
IMHO, HTTP, i.e., in the general sense, is not just for requesting webpages and resources for webpages.
I find HTTP/1.1 to be very useful. It is certainly not just for requesting webpages full of JS, CSS, images and the like. That is only one way I might use it. Perhaps HTTP/2 is the better choice for webpages. TBH, if using a "modern" graphical browser, I would be inclined to let it use HTTP/2. Most of the time I am not using a graphical browser.
I voted up because that is indeed neat tho.
I decided to try numbering the programs I write instead of naming them. I often use a prefix that can provide a hint.^1 For example, the yy prefix indicates it was created with flex and the nc in nc0 indicates it is a "wrapper script" for nc. If the program is one I use frequently, then I have no trouble remembering its number. In the event I forget a program number, I have a small text file that lists each yy program along with a short description of less than 35 chars.
1. But not always. I have some scripts that I use daily that are just a number. I also have a series of scripts that begin with "[", where the script [000 outputs a descriptive list of the scripts, [001, [002, etc. I am constantly experimenting, looking for easier, more pleasing short strings to type.
Each source file for a yy program is just a single .l file with a 3-char filename like 025.l, so searching through source code can be as simple as
grep whatever dir/???.l
If I put descriptions in C comments at top of each .l file I can do something like head -5 dir/???.l
Aesthetically, I like have a directory full of files with filenames that follow a consistent pattern and are of equal length. Look at the source code for k, ngn-k or kerf. When it comes to programming, IMO, smaller is better.