Google Exec Says It's A Good Idea: Open The Index And Speed Up The Internet
siliconvalleywatcher.com
siliconvalleywatcher.com
"After all, the value is not in the index it is in the analysis of that index."
The ability for a given search engine to innovate is based on having control of the index. The line between indexing and analysis isn't quite as clean as what is implied by the article, if only for the simple fact that you can only analyze what is in the index.
For example, at it's simplest and index is a list of what words are in what document on the web. But what if I want to give greater weight to words that are in document titles or headings? Then I need to somehow put that into the index.
What if I want to use the proximity between words to determine relevance of a result for a particular phrase? Need to get that info into the index, too.
In the end, what the author really wants is for someone to maintain a separate copy of the internet for bots. In order for someone to do that, they'd need to charge the bot owners, but the bot owners could just index your content for free, so why would they pay?
-Faster. You don't have the latency of millions of HTTP connections, but instead a single download. (Or a few dozen. Or a van full of hard drives.)
-Easier. The problem of crawling quickly but politely has been handled for you. The reading of sitemaps has been handled for you. The problem of deciding how deep to crawl, and when to write off a subsite as an endless black hole, has been handled for you. Etc.
-Predictable. Figuring out, in advance, how much it is going to cost you to crawl some/all of the web is, to say the least, tricky. Buying a copy with a known price tag provides a measure of certainty.
Of course, I am leaving out the potential pitfalls, but the point is there /are/ arguments in favor of buying a copy of the web (and then building your own index).
"The index" would just be a list of key/values. The key would be a URL, and the value would be the content located at that URL. There would also be some kind of metadata attached to the keys to indicate HTTP status code, HTTP header information, last-crawled date, and any other interesting data. From this data set, other, more appropriate indexes could be generated(for example via hadoop)
> In the end, what the author really wants is for someone to maintain a separate copy of the internet for bots.
Yes, but not for bots. It would be for algorithms.
> In order for someone to do that, they'd need to charge the bot owners
Probably.. The funding could be like ICANN, whose long-term funding comes from groups that benefit from its services.
> but the bot owners could just index your content for free, so why would they pay?
How would you create a copy of the internet for free? Are you just going to run your crawler on your home modem while you're at work? Where are you going to store all that data? How are you going to process it? How long is that going to take? Wouldn't it just be easier to(for example) mount a shared EBS volume in EC2 that has the latest, most up to date crawled internet available for your processing?
They have that index already. It's called "the internet". If you're storing the entire page content, that's not really much of an "index".
> it would be stupid store such a rich data set in such a dumb, information-losing format.
Discarding information is the entire purpose of indexing and doing things like map-reduce (emphasis on the reduce!): you discard worthless information so that the stuff you care about is of a manageable size.
The reason Google serves up search results quickly is because it carefully controls the data it has to wade through. Adding more data doesn't make it smarter, it makes it slower.
Fine. Let's call it an archive, or cache, or snapshot then. I only used that phrase because it was thrown around in the post I was responding to, and I specifically put that in quotes because I wasn't sure it was the right phrase to use.
> Discarding information is the entire purpose of indexing and doing things like map-reduce (emphasis on the reduce!)
No. Just no. An index is a data structure that is designed to quickly lookup its elements based on some key. It implies nothing about what your data(neither keys nor elements) look like.
Are you serious about "emphasis on reduce!"? "Reduce" is not named reduce because it inherently reduces the amount of data you are working with. It is called "reduce" because 1) an engineer at google liked the sound of it, and 2) it takes a list of intermediary values with all the same key. It is quite easy, and common, to have reduce functions which end up spitting out MORE data than you started with
> you discard worthless information so that the stuff you care about is of a manageable size.
Sometimes, but that doesn't work in this situation. How do you know what data is worthless before you know what algorithms will be applied to the data? The only acceptable solution is to keep a plaintext copy of the data retrieved from a particular URL.
So would they open an "index" of web page contents? In this case, why would another search engine access Google's "index" rather than the original server? The original server is guaranteed to be up to date, and there's no single point of failure.
I agree that Google's index is probably optimized to work with their search algorithm. From what the author claims, though, this doesn't mean that Google would be losing anything by allowing other engines to use the index, as "all the value is in the analysis" of the index.
1) You won't DOS the host site.
2) You don't have to respect robots.txt. If you need to crawl a 1 million page site, and its robots.txt restricts you to 1 page/second, then you'll have to wait for a long time. Downloading a crawl dump from a central repository would be much easier.
Yeah, consider a site like Hacker News, where the crawl delay is not 1 second but 30 seconds [1].
If you're trying to grab historical data like iHackerNews did [2], you might be better off hitting Google's webcache instead...except, to enumerate urls you have to scrape the search results pages, AFAIK.
This is why Google opening its webcache via an API is a GREAT idea.
This database would normally only accessible by search engines and the sites themselves could then disallow direct bot search in their robots.txt.
It occurs to me that this might have to be "invite only" - Google invites the sites they trust to put their data there but if they catch someone "cheating" in one way or another, they stop indexing. Plus they wouldn't have to invite really small sites.
It seems like this would be quite a complex project with a for the public good approach. Maybe it could work as an AWS project to sell amazon compute cycles.
As for your last two questions, there's nothing new whatsoever about them. Search engine pollution is ancient news.
Edit: looking at the other comments, doesn't seem to be that that many were able to get it. What a disappointing state of minds.
Centralizing it all so that one company has to run the index would mean that company would have to foot the entire bandwidth bill which they won't do, so it would end up distributed as the internet already is. It's just a shifting of the traffic to a centralized db which would need to be kept up to date.
There's more to an index than "hey, let's share an index". That's WHY we have multiple search engines: there will be differing views of what should be in that shared index, and how it should be constructed to facilitate varying needs. There are multiple search engines because each one plays off "hey, our index is better than theirs".
My snark makes the point: why not just have a single distributed/copied/shared library system, and do away with pricy competition? Answer: because that single system does not provide everything everyone wants, and people are willing to pay for a different selection & service.
The search engine and index are tightly coupled. They're huge, they're complex, they're expensive - and people think they can make a good buck by somehow doing it different. Create a "universal index", and someone will realize they can make money by making their own, leading right back to what we have today.
The TFA's key issue isn't really that multiple indexes are crawling his website, it's that they're gathering his data thru the most inefficient means possible - polling every page as often as practical. You don't want universal index, because short of banning competing indexes there will always be competing indexes. You want an agreed-on search interface: a means to serve the indexes what they want at a cost lower than what it costs them to poll-mine your site.
Yeah, a universal index is a bad idea. It ignores the fact (you'd think visitors to ycombinator of all sites would get this) that there is money in competition, ergo there will never be a single index. That, and it still doesn't solve the problem: it may reduce the number of crawlers sucking your bandwidth, but it's still polling every page as fast and often as it can.
Yeah, the more I read your post, the more I don't think you understand the point of the idea. The idea would be that many people could contribute to indexing the web, faster, and sharing that in a decentralized fashion. With open access to that index, anyone is able to innovate on that data and build their own search engine.
Having fewer crawlers would be a good thing and would indeed decrease the number of hits to your website. More importantly, the work of those crawlers could be distributed and thus page updates could happen even faster.
Even if you're right, that some how the indices themselves can be tuned and can be search-engine specific, who cares? That doesn't preclude or prevent an open index from existing.
BTW, I didn't really downvote anyone, nor did I have a problem with your comment. I just didn't understand it, and it was hot on the tails of another indecipherable comment. I think I understand your point now, but I feel we're either talking about two different things or thinking about them in vastly different manners.
It's not, it's called the web.
A rough distributed model could be implemented similar to the way we (hackers/coders) use github as a central repository for a distributed system. People contributing to the index on a private server could do whatever they want but since that instance of the index is not public, no one else will care about what the owner has done to it. Forks can be pushed to a public staging area where others can view it and verify it's accuracy, and then the major players can merge those changes into their forks.
The complaint (with github) that it is hard to figure out the canonical repo is also invalid in this model, as one can start with a fork of Google or Yahoo's public repo, and then build their own through merging or hacking directly on it, just like one can fork Linus' Linux kernel, and then merge in other's forks to incorporate other changes.
Remember, the index itself, as in the raw data taken in by GoogleBot or Yahoo! Slurp bot, would be the shared information. The analysis of the data, as in pagerank and other factors that Google decides makes one page more relevant to a keyword than the other, would not be shared as that is the bread and butter of each engine.
"I have a map of the United States... Actual size. It says, 'Scale: 1 mile = 1 mile.' I spent last summer folding it. I hardly ever unroll it. People ask me where I live, and I say, 'E6." - Steven Wright
I think the idea is good though. I think there would be a fight though to say who is the aggregator of the information. This would also mean whoever does distribute it has a stranglehold on the industry in terms of how and when it supplies this information.
I can see it's uses but I can equally see a lot of cons for the system not working or some serious amount of anti trust.
If you could get an unbiased 3rd party involved though and they built the database then I think that would work.
[1] http://www.amazon.com/Modern-Information-Retrieval-Ricardo-B...
I have read pretty big sections of Manning's Introduction to IR, and it served me fairly well as an introduction to the field. It's available online.[2]
[1] http://www.amazon.com/Modern-Information-Retrieval-Concepts-...
[2] http://nlp.stanford.edu/IR-book/information-retrieval-book.h...
Letting sites inject into the cache is an interesting idea, but Google will still have to spider periodically to ensure accuracy. Inevitably, a large number of sites will just screw it up, because the internet is mostly made of fail. This would leave Google with only bad options: If they delist all the sites to punish them, they leave a significant hole in their dataset. But if they don't punish them and just silently fix it by spidering, there is no longer any threat to keep the black hat SEOs in check. Either way, it would cause an explosion in support requirements and Google is apparently already terrible at that.
I don't suppose anyone has considered making an entry in robots.txt that says either:
last change was : <parsable date>
Or a URL list of the form
<relative_url> : <last change date>
There are a relatively small number of robots (a few 10's perhaps) which crawl your web site, all of the legit ones provide contact information either in the referrer header or on their web site. If you let them know you had adopted this approach then they could very efficiently not crawl your site.
That solves two problems;
@ web sites on the back end of ADSL lines but don't change often wouldn't have their bandwidth chewed by robots,
@ The search index would be up to date so if someone who needed to find you hit that search engine they would still find you.
I ran into this by chance when writing a wrapper to obfuscate e-mail addresses in mailing list archives. I didn't change the URL but had it served by a script instead of being a flat file. When it first went online, all of the robots kept crawling the files over and over. I finally made it supply the right mtime to Apache, which then did the right thing with the incoming IMS header, generating a 304 and not sending out new content.
It's possible this has regressed, but I would hope it hasn't.
Have you tried serving your pages with `Expires` and `Cache-Control` headers? I you give it - say - a timeout of a week, then a well-behaving client shouldn't retry before that time has went by.
Google invents lots of things, and I suspect that you are correct that they honor sitemap, they also abandon things they once thought were the answer (QR codes, Wave, Etc.). Since I have my own CMS for web pages I'll try adding this to it's output and see what GoogleBot does with it.
Not saying it's an accurate summary, but I don't think you were actually very curious!
1. If I were Microsoft, I wouldn't trust Google's index. How do I know they aren't doing subtle things to the index to give them an advantage?
2. Having the resources to keep a live snapshot of the web is one of the big players' advantages. Opening the index, while good for the web, would not necessarily be good for the company. Google could mitigate that by licensing the data: for data more than X hours old, you get free access; for data newer than that, you pay a license fee to Google. Furthermore, integrate the data with Google's cloud hosting to provide a way to trivially create map/reduce implementations that use the data.
3. On the other side, what a great opportunity the index could provide for startups. Maintaining a live index of the web is costly and getting more and more difficult as people lock down their robots.txt. Being able to immediately test your algorithms against the whole web would be a godsend for ensuring your algorithms work with the huge dataset and that your performance is sufficient.
Here's to hoping Google goes forward with it!
It's a cool idea though because Yahoo sucks up a ton of my bandwidth and delivers very little in SEO traffic. On most of my sites now I have a Yahoo bot specific Crawl-Delay in robots.txt of 60 seconds, which pretty much bans them.
And let the site notify the indexer when things change, so all the bandwidth isn't used looking for what's changed. Actual changes could make it in to the index more quickly if the site could draw attention to it immediately rather than an army of robots having to invade as frequently as inhumanly possible. The selected indexer could still visit once a day or week to make sure nothing gets missed.
Their attitude is to take everything in but not to let you automate searches to get data out.
This is the biggest problem I have with search engines - you want to deep index all my sites? Fine, but you better let me search in return - deeper than 1000 results (and ten pages). Give us RSS, etc.
You would write something that takes the information they are referring to in this article, it's how you digest and index that information yourself that makes the difference
For _most_ of the websites, its in _their_ interest to have good SERP ranking. Not the other way around.
"The index"? Feature extraction is the most complex part of almost any machine learning algorithm, and search is no different. Indexing full text documents is a really difficult task, especially if you take inflected languages into account (English is particularly easy).
I don't see a way to "open the index" without disclosing and publishing a huge amount of highly complex code, that also makes use of either large dictionaries, or huge amounts of statistical information. It's not like you can just write a quick spec of "the index" and put it up on github.
FWIW, I run a startup that wrote a search engine for e-commerce (search as a service).
That's what the author means by 'index'.
Edit: SmugMug seems to fall into this category: http://don.blogs.smugmug.com/2010/07/15/great-idea-google-sh...
Also interesting:
And if you think about it, the robots are much harder to optimize for – they’re crawling the long tail, which totally annihilates your caching layers. Humans are much easier to predict and optimize for.
I think a more correct title for this post would have been "Google exec tries to politely dismiss a silly idea", but of course that's not as punchy.
If it's the former all this really does is move the burden from sites to Google, and introduces a single point of failure.
If it's the latter, which seems unlikely, what incentive does Google have to share that data? It's part of their competitive advantage.
A single public index would expose this data to stronger analysis (or even plain reading), not just Google search queries.
If you think about it, it does make sense in a lot of respects. I have dealt with a lot of companies that sell data, the only difference is this data is freely available to everyone so everyone thinks they should crawl the information themselves.
The only people who lose out are the people paying the bandwidth bills. The internet would actually be slower due to the amount of information passing around when it is not needed.
This idea makes more sense the more we discuss it
Bandwidth is not nearly as expensive as the overhead of such search index centralization.
This won't be accepted. And even legally, this is not possible due to copyright laws in different countries.
If it's the former, this seems like it wouldn't help speed at all.
EDIT: To clarify, how does reducing the amount of bandwidth used speed up anything? Why am I being downvoted for this?
Of course, if this is the primary issue, I don't see why Google couldn't just implement a closed push-based index. When your site updates, you push the changes to Google. The index is still closed but it solves the bandwidth problem without opening Google resources.
Considering there are dozens of bots that will crawl one's site in "random" order, they all begin to wreak havoc on caches, as each robot and the humans don't browse in uniform.
People who downvoted you didn't understood or think your comment is not pertinent.
It's not literally a speed comparison, but you can imagine how if you could optimize away all the robot requests, you could do something else with the time and resources you'd formerly spent serving them.
The smugmug addendum to the story explains this pretty clearly.
Problem solved... like a million internet years ago.
Isn't the proposal clear enough?
1. Optimize the indexing process so that we avoid each search engine crawling independently every site.
2. Devise a method to refresh the index when the content changes (hash, date...)
Seems resonable enough, to me.