Source: I run a think tank focused on Google’s web crawling advantage and have been studying stuff like this for a couple years now.
Source: I run a think tank focused on Google’s web crawling advantage and have been studying stuff like this for a couple years now.
That does seem like a good idea since it amounts essentially to a database of stuff that Google does not own (by construction) and is of public utility.
If you would like to read more about all this, please checkout https://knuckleheads.club
Some feedback:
- A cost of membership runs contrary to establishing this group, especially at such a high recurring charge.
- I'm not sure what your software/AWS situation looks like, but 20 million robots.txt files acquired from Common Crawl is something I can analyze on my PC. It doesn't seem to presently justify such high costs.
- Prioritize building a mockup index with an intuitive frontend. This is essential for non-technical people to understand
- Exclusively talk with EU legislators (they are motivated, whereas nothing will happen in the US).
I think the price for membership dues is reasonable and many people agree evidenced by them signing up. I think I might start a petition that is free to sign up on though, thank you for the inspiration!
It is possible to analyze those files on the pc, it just takes a much longer time. The analysis is an iterative process and so the faster the computers the faster the iterations and process go. I was analyzing them on my pc with python for the first year until it got too slow and my I am using an aws server with some rust and that is going much better. I also need to increase the number of files analyzed by about two orders of magnitude soon as well.
Great idea, very cool. That’s going on the todo list!
And I am going to be reaching out to and speaking with whoever is interested. One of the fun things about this is that it is an international dynamic, with some jurisdictions having abilities that others don’t. For example, the UK CMA has subpoena powers that the US Congress lacks and got a ton of information out of Google and Bing that shocked me. The US has the ability to get the CEO’s to show up to hearings while the UK does not in the same way. Why limit ourselves to one government when there are so many to mix and match from here?
??
The first is that Google gets much more access to pages on websites than everybody else. You can see this by examining the robots.txt files of various websites[0]. I've been doing this for several years now and Google has a consistent advantage across many thousands websites that I've looked at. This adds up to a significant advatnage and many search engine operators complain about how it hampers their ability to compete with Google[1].
The second is that Google gets to ignore crawl delay directive in robots.txt while other search engines don't[2]. Website operators cannot tell Google how fast they want their website crawled, they can only request that Google slow down. If another search engine tried to do what Google does, they would likely be blocked by many important websites.
If you would like to read more about this, please checkout https://knuckleheads.club/
[0] https://pdf.sciencedirectassets.com/robots.txt
[1] https://www.nytimes.com/2020/12/14/technology/how-google-dom...
[2] https://www.seroundtable.com/google-noindex-in-robots-txt-de...
For me it was a problem of having lots of pages, and having a high cost per request (due to the type of website it was).
For other websites, it is not necessarily about the volume of traffic from bots, but the risk of web scrapers getting their proprietary data. They're fine with Google scraping their info because that's where their traffic comes from. They're not okay with some random bot scraping them because it could be taking their content and republishing it, or scraping user profile data, or using it for some nefarious/competitive purpose.
That's some weird logic, to me at least. That data is literally given away to everyone but some people or organizations can't have it? If you want to control access to it, maybe at least require people to register before they can see it? Is it even proprietary if it's public with no access control whatsoever?
This for-profit internet is just really such a parallel universe to me.
It’s a different world where there are no laws or prices or contracts really.
I know I have been a contrary commentor in this thread, but I hear you with this. What a monster we have built, and what always gets me is how trivial everything is. So much capital is flowing through these ephemeral software systems that, if gone tomorrow, would be ultimately inconsequential to mankind.
> and what always gets me is how trivial everything is
Whenever I read about corporations and how they work, I always inevitably ask myself the question "where the hell does enough work to keep this many people busy even come from". Everything is ridiculously overengineered to meet imaginary deadlines.
It's often a question of quantity. LinkedIn probably doesn't care about you scraping a few profiles, but if you're harvesting every bit of their publicly-available data, then they get a little scared that you're building something that's going to compete with them.
Same with Instagram, or Facebook, for example. Though in this case it's probably more of a user-privacy issue - at least that's what they say.
It's not really weird logic to me - seems to make sense.
> If you want to control access to it, maybe at least require people to register
Most of the time they can't do this because they need the Google traffic. LinkedIn wants a result in the SERP for Bob Smith when you search for "Bob Smith" because that helps them get signups. Google won't list the page if that content is gated by a sign-in/register page.
How would that work? I'm pretty curious about this if there's anything out there to read.
EDIT: Ah I saw the link to Knuckleheads' Club :P
What incentive would Google have to continue populating that index?
Would I be breaking the law if I independently crawled and hosted an index without publishing an API for it?
> since no one other than Google is really allowed to crawl the web
Maybe this is the problem that needs solving.
Presumably they would still want to run google.com and make money off of it.
> Would I be breaking the law if I independently crawled and hosted an index without publishing an API for it?
No. You would not get the advantage that Google gets when it crawls the web and so would not have access to a large amount of data that nobody else has access to.
Updated based on edit of parent post:
> Maybe this is the problem that needs solving.
Why have websites waste the money to serve all those requests all over again? Why don't we have Google share the results and we can use that money to do more productive things than recreating that work? I don't think website operators would be happy if there were a hundred more crawlers out there crawling as much as Google does now.
Perhaps other search engines should spoof GoogleBot. Browsers have being doing that since forever spoofing Netscape (Mozilla), Safari, etc. for the same reason.
> Why don't we have Google share the results and we can use that money to do more productive things than recreating that work?
This sounds like a common fallacy of people criticizing the free market. Duplicated effort looks wasteful but turns out to be far more productive than the lack of incentive that comes with not being able to profit from your work/investment.
Many website operators do actually block crawlers from non Google search engines and it's because the cost of being crawled isn't worth it to them. Here's a good quote from one such webmaster:
As a webmaster I get a bit tired of constantly having to deal with the startup crawler du jour.
From law firms looking for DMCA violations to verticals search engines, to image aggregators, to company intelligence resellers… It feels to me that everybody and their brother has gotten into spidering sites.
With 10,000s of pages that have content that is only relevant to a targeted audience who is perfectly able to find us on the majors, I do not hesitate to block (and possibly ban) when I see an aggressive crawler that does not provide me or my customers with direct benefits.
Taken from http://www.skrenta.com/2008/04/cuill_is_banned_on_10000_site...> Perhaps other search engines should spoof GoogleBot. Browsers have being doing that since forever spoofing Netscape (Mozilla), Safari, etc. for the same reason.
People have tried this and it doesn't work. Google provides ways to check to make sure traffic is coming from Google IP addresses and practitioners and academics study how to spot fake Googlebots. https://developers.google.com/search/docs/advanced/crawling/... https://blogs.akamai.com/2014/07/search-engine-impersonation... https://ieeexplore.ieee.org/document/8421894
> This sounds like a common fallacy of people criticizing the free market.
I am asserting that crawling the web is a natural monopoly. This means that the free market has failed and that it is not possible for the market to heal itself in this regard. There is significant evidence that this is the case and I imagine you'll be hearing more and more about it soon.
The benefit varies with the quality of the search engines, and that will vary between search engines, but it does get larger the more a search engine is used, so a cost/benefits analysis may show Google and a few other large ones are the only ones worth supporting.
Spoofing crawler identity completely defeats the point of the honor-system robots.txt.
Only running the right SEO wasn't sufficient then and I certain wish that criteria remained.
Of course, "quality" applied is a broad term indeed. But there no being a single certain measure of quality doesn't mean there aren't some things that most people would call crap. Google employs a large number of search quality raters and these folks can likely distinguish the terrible from the OK.
Google is far from ideal and getting worse but you seem to be exposing a pure democracy of algorithm conformance, which is guaranteed to result in hot, steaming idiocy.
The discussion is about Pininterest in text searches. My only objection to Pinterest in image searches is that it's paywalled/login-walled.
But larger issue is I'd acknowledge some search results just aren't what I in particular want but I'd claim other results are actually objectively low quality and you seem to want to push things to realm of pure subjectivity, any old crap is something someone might want and who I am to deny to them that pnis enlargement pill.
I would say that the results Google returns involve a number of factor/filters. A. What's considered mainstream, what appeals to many B. What's could be more or less objectively called quality. C. SEO, What results just slip through based on the page spending a lot of time and money appearing like A or B to the algorithm (but not being that).
And Google spends a huge amount of time and money trying to keep C from being the only thing BUT that's still not enough because there's a lot of time and money spent on the other end. At the time, this tug-of-war serves as a moat keeping competitors out. Any Google competitor would have to invest similar amounts of money.
Google has, does and always will make choices about what is best and that the definition of best is up to their discretion.*
I'd love to have a decent Google competitor. But I don't think you have won many friend here by implying that Google's discretion is just arbitrary (as a number of your posts here seem to imply to me). Google won, back when they had competition, by caring a lot about the, uh, quality of their search results. Now that they've won, they're slipping into other things and moreover, the SEO trash are nipping their heels.
Which is to say you won't build Google alternatives on "democracy" but on some concept of quality people will want.
Continuing out the logic of your suggestion here, you would seem to imply that every other currently existing search engine besides Google also has bad quality as well in this regard. I say this because if other search engines had better quality that Google, they would be doing better than Google as per your suggestion. But, obviously Google is on top in a big way, so I have to ask, do you think the entire search engine industry, not just Google, has poor quality search results? What do you know that they don't?
Me earlier: larger issue is I'd acknowledge some search results just aren't what I in particular want but I'd claim other results are actually objectively low quality
It seems like you just a kind of "shtick". Anyone talking about quality, no how much nuance they add to it, gets thrown the same "how do you your ideas of quality are right". I already mentioned that quality is piece of Google formula and Google spend real money attaining that, employs thousands of people to rate search quality. Of course I'm not claiming to be a personal expert on quality. You can read further what I actual wrote above.
Making that historical engagement data public is not feasible/realistic, IMO.
I have never worked for a search engine company (or devoted much time to SEO), but I would suggest the click-and-query data advantage is invaluable to cementing Google's supremacy over any newcomers.
I’ve heard search engine operators complain about the index consistently while they seem spilt about the click and query data. For example, DuckDuckGo and StartPage have made pretty good businesses out of pointedly not collecting click and query data. For my part, I think it’s the dual lock on distribution via exclusive agreements on being the default search engine and the advantages it has when web crawling that cements Google’s dominance over the search engine market.
I've only ever met one person that had anything to do with a think tank at a party about a decade ago, and I never got her contact details so I'm still very much in the dark.