There’s not much room to squeeze in when your competitors hold the keys to 15 million top websites.
I find it wild that "at scale" we can bypass anti-bot measures, but just "normal" internet use (i.e Non-Google Browser or VPN) will throw a million captchas at you.
cgnat is pretty bad too.
The big issue you sidestep not at scale is you can come from a single, residential IP with a good reputation.
Mandatory captchas for simply viewing a page are rare - most are saved for high impact actions like account creation.
When this does happen for a simple page view, AI is extremely good at solving basic captchas - especially basic “click the box” captchas.
If you don’t want to pay for AI, there are decaptcha services where someone in Southeast Asia solves the captcha for fractions of a penny. Save the cookies after a successful solve and you’re probably good in the future.
If you don’t want to pay for someone to solve a simple check the box captcha a little bit of attention and some properly simulated clicking (IE not a JavaScript injected event) will often work. Just don’t click literally the exact middle, fuzz the coordinates and you’re good.
Why would website authors _want_ to prevent crawling by other search engines?
Previously captcha was just for spam limiting, but I actually looked at our system logs and about half of traffic was bad behaving scrapers.
In logs I see these scrapers are hitting every link on the page. If you have a collection page then it's hitting every filter option and then hitting each pagination button, the different sort orders, etc. People running something like Forgejo it will hit every commit.
If you have expensive to compute pages, they're getting hit by these incredibly naive bots that don't respect any robots.txt or discriminate on what they do.
So if another search engine does arise, it won't find anything useful, because the useful content on the web has been buried under slop, and largely removed. Your best bet today is a curated directory, sorta like the original Yahoo, where you allowlist the web to only real sites, download them, and make them searchable. I think this is actually Kagi's approach. But the open web as we knew and loved it is dead.
https://blogs.microsoft.com/blog/2023/02/07/reinventing-sear...
[1]: https://alternativeto.net/software/google-search/?license=co...
Very few of the smaller search engines actually do their own indexing for exactly this reason.
At this point with Google contributing so much to the Trump administration, I'm not sure which is worse.
- Must not alter the order of Google's search results - Must not alter the appearance or placement of Google-inserted ads
When I use google, usually from my phone, I am reminded of why I don't use google on desktop.
With the announcement of this move by them, I just manually removed google as an address bar search engine option in all my browsers on desktop and mobile.
Human produced content should be separated from sites primarily hosting slop. That seems solvable?