Quora Blocks Startup Search Engines
readwriteweb.com
readwriteweb.com
Our project was a bit odd, in that we had millions of documents available to browse. We wanted them all to be available to search engines, but it was an operational headache to keep that much data moving.
Our normal daily traffic might have been 10,000 users with perhaps 50K page requests. We built our first small web server to handle this load.
We were caught blindsided by the Googlebot, which from day one demanded hundreds of thousands of pages per day. Googlebot offers no option to slow it down, and if you try to slow it down on purpose your search ranking will suffer.
So 90% of our early operational burden was devoted to keeping Googlebot happy. This was an unexpected burden, but it was worth it to be in the Google index, because they sent a lot of traffic our way and kept our fledging startup alive.
But it was definitely not worth it to spend equivalent operational expense to be in marginal search engines that don't refer much traffic. To a site operator, they don't offer any value in exchange for what they demand.
It was very interesting to see the relative volume of requests that we got from Googlebot versus the other crawlers. While Googlebot found us on day one, it took over 3 months before we appeared in Microsoft's index, and even then, they had perhaps only 10% of the coverage of our site when compared to Google.
http://www.google.com/support/webmasters/bin/answer.py?answe...
What if Googlebot starts at a non-root location (via a link from another site)?
Will it always start at root?
I'm not saying you were in the wrong, but that is the stance I would (and do, with my current project) take. And of course, when we are ready, we will open the door to all indexing.
I know at minimum GoogleBot and Yahoo! Slurp respect this value. I believe Google AdsBot (Landing page URL verification for AdWords) also supports this value, but GoogleBot and AdsBot each will crawl at the rate you specify (so if you specify a Crawl-Delay of 1, each bot will crawl at ~1qps). I don't know if it is in the spec, but fractional crawl delays appeared to be respected (Crawl-Delay: 0.25 would result in ~4qps, for example).
I too have fought with this problem - at times more than 90% of the capacity of the site I was running was devoted to serving up content for bots. I sympathize with small (and large) companies who don't want to add capacity so that new/random bot can add another <x> QPS to the daily baseline load.
The first two lines are:
# If you operate a search engine and would like to crawl Quora, please
# email info@quora.com. Thanks.
Looks sensible to me.Five bots are listed. Perhaps they're the only one who bothered to email. Or perhaps not.
There is no rule in there that denys other search engines, just some rules that set some slightly different rules for different ones.
EDIT: I see now -therule at the end is rather specific, and wouldn't allow crawling the entire site. Still - without actual complaints from someone who contacted them and got no response, I see no evil.
EDIT: my bad. I got confused between Facebook and Quora
Disallow: /
means?It seemed self-explanatory here.
I was being perhaps a bit sarcastic...
Quora should just get over themselves.
The same goes for life off of the Internet, which of course could be recorded or reported, and subsequently placed on the Internet.
The real problem is people mistakenly believe they can secretly operate with impunity in an age of increasing transparency.
I think the search engine privacy feature is a fantastic innovation from Quora to help absolve potential problems in the future. (eg: you google my name and see my Quora answers on Sex)
If you wrote public answers on questions of sex, why should they be hidden? --You made your answers public intentionally and also provided the answers intentionally so others would hopefully learn from them. Additionally, you did both while knowing your name is attached to them. In other words, you want to be known both for and by your contributions.
If you did not wish to be known for and by your public contributions, then the only answer is to either not make public contributions, or go to rather extreme lengths to only make public contributions anonymously. The latter fraught with plenty of caveats to maintain anonymity correctly, so the only real answer is to only make public contributions you would want attributed to you.
Back in the 80's when Chico State was the number 1 party school in the US, recruiters would show up in droves looking to get new employees. The unstated reason was very simple; if someone could graduate from Chico, it pretty much proved the person could not only get things done but also have a really fun time doing it.
Who would you rather work with? --A fiercely academic but socially limited graduate, or the guy who can get stuff done and also do keg stands?
If I searched on your name and found your public opinions, I'd think far more of you for having the stones to make your own contributions than I would be concerned about your predilection for midget porn.
Of course, the trouble with trying to be non-judgmental and unbiased is assuming others will do the same.
http://www.google.com/#sclient=psy&hl=en&q=site:duck...
Way to play the openness card Gabriel.
Wait, since when was Blekko big? Also, why has every article about search engines recently ended up talking about DDG? :p
http://www.readwriteweb.com/hack/2010/12/the-secrets-behind-...
That's pretty big in my book.
Because they're the new darling, and tech types like them. It's always nice to see someone come in as a substantial underdog and make a real effort to succeed. Also, their results seem to be at least as good as Google's which is pretty impressive.
They're Bing with some twiddles on top. DDG uses Yahoo! BOSS, which is Yahoo's API to its search service, which is now powered by Bing. Your DDG results all originally come from Bing's index, they just get filtered and possibly re-ranked by some custom code Gabriel's written.
How do you get your results?
From over 30 sources, including DuckDuckBot (our own crawler), crowd-sourced sites, Yahoo! BOSS, embed.ly, WolframAlpha, EntireWeb, Bing & Blekko.
Just do some searches and compare to Bing, etc., e.g. http://duckduckgo.com/?q=hacker+news vs http://www.bing.com/search?q=hacker%20news vs http://www.google.com/search?hl=en&q=hacker%20news
>We never stopped crawling; we just scaled it back for specific purposes, mainly now for spam detection and zero-click info
So if you've scaled back the crawling and are mostly using it just for spam detection and zero-click info (forgive my ignorance, but I'm not sure what zero-click info is) then does that mean that you are getting the majority of your result data from outside APIs like BOSS?
Edit: And does that mean that you store those pages in your database and use your own ranking algorithm, or are you using their ranked results for a query and then rearranging them on the fly?
About 50% of queries show information from my index. I've concentrated on the fat head of the search space, where I get a lot of volume (1 and 2 word queries). For long-tail stuff, yes, the majority of results come from external APIs, but even there it is inaccurate to say they are all derived from a direct call to one API. It really varies by query. Some will look like Bing for sure, but others will look very different. I'm not going to disclose everything I do of course.
Well, I didn't even think of it that way. Naturally you're probably looking at these types of criticisms under a magnifying glass since it's your project.
>About 50% of queries show information from my index. I've concentrated on the fat head of the search space, where I get a lot of volume (1 and 2 word queries).
See now that's actually good to know! From many people's point of view (including mine), DDG is Bing with some twiddles on top. Maybe out of ignorance, but what can you expect? If you aren't already enamored by the idea of DDG then you probably don't care enough to find out what you're really doing on the back-end, which is what really matters most to me in these discussions. (Not the results since unless I'm told otherwise I just assume it's Bing or Yahoo.)
Edit: It's quite humorous to be down-voted by the fan boys for expressing an honest and IMO helpful opinion to the owner of DDG. If all he gets is blind love he won't know how to get the rest of us on board. Believe it or not, the down vote button isn't the disagree button. If you have something constructive to add, say it.
I'm not going to downvote you, but next time just say what you have to say. If people like it and upvote it, great. if not, it's a meaningless metric on a fairly small social news site. There are far more important things to worry about.
To add to that, if I were trolling or being rude or vicious I would not have edited to bring up the down-voting. I hate to see this site slip into the bad habit of down-voting legitimate and helpful points due to simple disagreement with the commenter. With so many new users coming to HN daily it's important to occasionally remind people that this isn't Reddit, it's not Digg, and it's not Slashdot.
My impression from the FAQ and your comments here is that you start with Yahoo! BOSS results, merge in other results that you suspect may be useful from your own crawl, filter out known spam pages, and possibly re-rank things afterwards according to whatever algorithms you have. Hence, "Bing with a bunch of twiddles on top." I don't really get a different impression after reading this comment - "50% of queries show information from my index" could mean as little as reranking one result, and query distribution is known to be heavily biased towards the fat head.
Although, I just looked at one not-quite-head-but-not-quite-long-tail query ([modern family episode list]):
http://duckduckgo.com/?q=modern+family+episode+list
http://www.bing.com/search?q=modern+family+episode+list
And the top 3 are identical, but the rest of the top 10 are about as different from Bing as Bing is from Google. Twiddles can be fairly powerful.
You obviously don't have to disclose everything you do in response to criticism - Google and Bing certainly don't. But that was my perception from your FAQ and posts, and certainly seems to be the perception of many other HN users as well.
I have looked at those search APIs and they don't let you write your own ranking formula. They just give you back 10-20-30 top results. There is no info on term frequency, inverted document frequency, pagerank or any other component of a decent ranking formula. So you can rank yourself the pages from your own crawls, but you can't merge this nicely with the API results.
If I were Garbiel, I would merge them in some crude way basically assuming my results are better, and add those URLs to the list for my crawler. The top URLs for all one and two word queries aren't that many.
Those are all big wins by themselves. And speaking as a user, if it's doing all three I am only mildy curious about how it gets those results.
http://www.paulgraham.com/submarine.html
Your job as a startup is to sit around all day coming up with seemingly important news stories that just happen to revolve around your company. Get some press contacts to feed those stories to, and you're sorted for marketing.
Bonus points for attacking the Big Competitor in a public way, thus forcing major news outlets to cover the "story" and ask Big Competitor for a quote about you.
"[Quora's] robots.txt file explicitly grants Google, Bing, Blekko and other big players access"
Blekko is a big player? It launched two months ago and it's still in public beta. According to Compete.com [1], it attracts only 120,000 visitors a month.
I'd say that if Blekko is a big player, DuckDuckGo is one too.
http://www.readwriteweb.com/hack/2010/12/the-secrets-behind-...
Has DuckDuckGo published the details of their data-center anywhere?
Also, Blekko may have launched recently, but it was founded in mid-2007.
Also, when disallowing a spider you can be sure that you won't get ANY visitors from them, as you are not in their index to start with.
I mean: if I disallow any bot but Googlebot, I will have 100% Google referrals by design.
I can understand that in an emergency situation you might want to block some bots, but blocking all of them because they are currently small feels a) a bit unfair due to not giving them the chance to grow and b) shortsighted due to never knowing if they might be getting big enough or the visitors they deliver might be better ad-clickers
If a site really doesn't want you there, they can block you at the IP/User-Agent level, which is what Quora will end up doing.
Essentially, the existence of robots.txt and meta noindex has made courts more comfortable ruling that the index of search engines constitutes fair use.
Interesting link, however. I was unaware of that case.
1. Honoring robots.txt is optional, so this doesn't actually offer any form of asset protection.
2. They are losing serious amounts of exposure that they would reap if they'd lift this restriction so that smaller engines can find them.
thanks!
5.3 You agree not to access (or attempt to access) any of the Services by any means other than through the interface that is provided by Google, unless you have been specifically allowed to do so in a separate agreement with Google. You specifically agree not to access (or attempt to access) any of the Services through any automated means (including use of scripts or web crawlers) and shall ensure that you comply with the instructions set out in any robots.txt file present on the Services.
here ya go: http://lmgtfy.com/?q=how+to+be+helpful
They do allow automated queries through their APIs for at least some of their search verticals.
It looks to me like it doesn't block anything - it does have some specific settings tailored for various big-boy search engines, with a catch-all rule at the end for those not otherwise defined.
It also says right there at the top "# If you operate a search engine and would like to crawl Quora, please # email info@quora.com. Thanks. "
So they're trying to get the right data to the right search engines the right way - makes sense to me.
User-agent: * Allow: /$ Allow: /about$ Allow: /about/ Allow: /jobs$ Allow: /challenges$ Allow: /press$ Allow: /login$ Allow: /login/ Allow: /signup$
Craiglist policy of specifically disallowing only classified search engines should look kinder in retrospect.
It is reasonable to block sites which aim to merely re-frame your content within a "search" parameter. But that kind of thing can be dealt with terms of service.
The problem is that it takes time to actively filter against offenders.