That does seem like a good idea since it amounts essentially to a database of stuff that Google does not own (by construction) and is of public utility.
If you would like to read more about all this, please checkout https://knuckleheads.club
Some feedback:
- A cost of membership runs contrary to establishing this group, especially at such a high recurring charge.
- I'm not sure what your software/AWS situation looks like, but 20 million robots.txt files acquired from Common Crawl is something I can analyze on my PC. It doesn't seem to presently justify such high costs.
- Prioritize building a mockup index with an intuitive frontend. This is essential for non-technical people to understand
- Exclusively talk with EU legislators (they are motivated, whereas nothing will happen in the US).
I think the price for membership dues is reasonable and many people agree evidenced by them signing up. I think I might start a petition that is free to sign up on though, thank you for the inspiration!
It is possible to analyze those files on the pc, it just takes a much longer time. The analysis is an iterative process and so the faster the computers the faster the iterations and process go. I was analyzing them on my pc with python for the first year until it got too slow and my I am using an aws server with some rust and that is going much better. I also need to increase the number of files analyzed by about two orders of magnitude soon as well.
Great idea, very cool. That’s going on the todo list!
And I am going to be reaching out to and speaking with whoever is interested. One of the fun things about this is that it is an international dynamic, with some jurisdictions having abilities that others don’t. For example, the UK CMA has subpoena powers that the US Congress lacks and got a ton of information out of Google and Bing that shocked me. The US has the ability to get the CEO’s to show up to hearings while the UK does not in the same way. Why limit ourselves to one government when there are so many to mix and match from here?
??
The first is that Google gets much more access to pages on websites than everybody else. You can see this by examining the robots.txt files of various websites[0]. I've been doing this for several years now and Google has a consistent advantage across many thousands websites that I've looked at. This adds up to a significant advatnage and many search engine operators complain about how it hampers their ability to compete with Google[1].
The second is that Google gets to ignore crawl delay directive in robots.txt while other search engines don't[2]. Website operators cannot tell Google how fast they want their website crawled, they can only request that Google slow down. If another search engine tried to do what Google does, they would likely be blocked by many important websites.
If you would like to read more about this, please checkout https://knuckleheads.club/
[0] https://pdf.sciencedirectassets.com/robots.txt
[1] https://www.nytimes.com/2020/12/14/technology/how-google-dom...
[2] https://www.seroundtable.com/google-noindex-in-robots-txt-de...
For me it was a problem of having lots of pages, and having a high cost per request (due to the type of website it was).
For other websites, it is not necessarily about the volume of traffic from bots, but the risk of web scrapers getting their proprietary data. They're fine with Google scraping their info because that's where their traffic comes from. They're not okay with some random bot scraping them because it could be taking their content and republishing it, or scraping user profile data, or using it for some nefarious/competitive purpose.
That's some weird logic, to me at least. That data is literally given away to everyone but some people or organizations can't have it? If you want to control access to it, maybe at least require people to register before they can see it? Is it even proprietary if it's public with no access control whatsoever?
This for-profit internet is just really such a parallel universe to me.
It’s a different world where there are no laws or prices or contracts really.
I know I have been a contrary commentor in this thread, but I hear you with this. What a monster we have built, and what always gets me is how trivial everything is. So much capital is flowing through these ephemeral software systems that, if gone tomorrow, would be ultimately inconsequential to mankind.
> and what always gets me is how trivial everything is
Whenever I read about corporations and how they work, I always inevitably ask myself the question "where the hell does enough work to keep this many people busy even come from". Everything is ridiculously overengineered to meet imaginary deadlines.
It's often a question of quantity. LinkedIn probably doesn't care about you scraping a few profiles, but if you're harvesting every bit of their publicly-available data, then they get a little scared that you're building something that's going to compete with them.
Same with Instagram, or Facebook, for example. Though in this case it's probably more of a user-privacy issue - at least that's what they say.
It's not really weird logic to me - seems to make sense.
> If you want to control access to it, maybe at least require people to register
Most of the time they can't do this because they need the Google traffic. LinkedIn wants a result in the SERP for Bob Smith when you search for "Bob Smith" because that helps them get signups. Google won't list the page if that content is gated by a sign-in/register page.
How would that work? I'm pretty curious about this if there's anything out there to read.
EDIT: Ah I saw the link to Knuckleheads' Club :P
What incentive would Google have to continue populating that index?
Would I be breaking the law if I independently crawled and hosted an index without publishing an API for it?
> since no one other than Google is really allowed to crawl the web
Maybe this is the problem that needs solving.
Presumably they would still want to run google.com and make money off of it.
> Would I be breaking the law if I independently crawled and hosted an index without publishing an API for it?
No. You would not get the advantage that Google gets when it crawls the web and so would not have access to a large amount of data that nobody else has access to.
Updated based on edit of parent post:
> Maybe this is the problem that needs solving.
Why have websites waste the money to serve all those requests all over again? Why don't we have Google share the results and we can use that money to do more productive things than recreating that work? I don't think website operators would be happy if there were a hundred more crawlers out there crawling as much as Google does now.
Perhaps other search engines should spoof GoogleBot. Browsers have being doing that since forever spoofing Netscape (Mozilla), Safari, etc. for the same reason.
> Why don't we have Google share the results and we can use that money to do more productive things than recreating that work?
This sounds like a common fallacy of people criticizing the free market. Duplicated effort looks wasteful but turns out to be far more productive than the lack of incentive that comes with not being able to profit from your work/investment.
The benefit varies with the quality of the search engines, and that will vary between search engines, but it does get larger the more a search engine is used, so a cost/benefits analysis may show Google and a few other large ones are the only ones worth supporting.
Spoofing crawler identity completely defeats the point of the honor-system robots.txt.
Many website operators do actually block crawlers from non Google search engines and it's because the cost of being crawled isn't worth it to them. Here's a good quote from one such webmaster:
As a webmaster I get a bit tired of constantly having to deal with the startup crawler du jour.
From law firms looking for DMCA violations to verticals search engines, to image aggregators, to company intelligence resellers… It feels to me that everybody and their brother has gotten into spidering sites.
With 10,000s of pages that have content that is only relevant to a targeted audience who is perfectly able to find us on the majors, I do not hesitate to block (and possibly ban) when I see an aggressive crawler that does not provide me or my customers with direct benefits.
Taken from http://www.skrenta.com/2008/04/cuill_is_banned_on_10000_site...> Perhaps other search engines should spoof GoogleBot. Browsers have being doing that since forever spoofing Netscape (Mozilla), Safari, etc. for the same reason.
People have tried this and it doesn't work. Google provides ways to check to make sure traffic is coming from Google IP addresses and practitioners and academics study how to spot fake Googlebots. https://developers.google.com/search/docs/advanced/crawling/... https://blogs.akamai.com/2014/07/search-engine-impersonation... https://ieeexplore.ieee.org/document/8421894
> This sounds like a common fallacy of people criticizing the free market.
I am asserting that crawling the web is a natural monopoly. This means that the free market has failed and that it is not possible for the market to heal itself in this regard. There is significant evidence that this is the case and I imagine you'll be hearing more and more about it soon.