"How do I set the User-Agent string in Java?" - L. Page (1996)
guyro.typepad.com
guyro.typepad.com
It's blisteringly fast, low CPU usage, and would suit the core task of crawling websites extremely well. The async NIO libs are fantastic for network io.
What would you use and why? (And why would you not use java)
> Extremely well suited to backend tasks running on servers for months on end without crashing.
Running for months without crashing - maybe. But by the end of that month (well, week really) it will be so slow (i.e. due to memory leaks ironically) that your only option will be autorestarting it every now and then...
AFAIK Larry and Sergey chose Perl at the beginning. Now it should be mostly Python and C.
You can certainly run for months without issue (memory/crash/speed) as long as you don't have any leaks in your own code.
I'm pretty sure Java is still widely used at Google.
If I was writing the google crawler from scratch today, I'd certainly start with Java, then probably use perl/python for less critical scripting glue, and maybe rewrite any CPU intensive stuff in C/asm.
Currently, the predominant business model for commercial search engines is advertising. The goals of the advertising business model do not always correspond to providing quality search to users. For example, in our prototype search engine one of the top results for cellular phone is "The Effect of Cellular Phone Use Upon Driver Attention", a study which explains in great detail the distractions and risk associated with conversing on a cell phone while driving. This search result came up first because of its high importance as judged by the PageRank algorithm, an approximation of citation importance on the web [Page, 98]. It is clear that a search engine which was taking money for showing cellular phone ads would have difficulty justifying the page that our system returned to its paying advertisers. For this type of reason and historical experience with other media [Bagdikian 83], we expect that advertising funded search engines will be inherently biased towards the advertisers and away from the needs of the consumers.
Looks like they solved the problem by turning it on its head.
Write down the problem.
Think real hard.
Write down the solution.
The Feynman algorithm was facetiously suggested by Murray Gell-Mann, a colleague of Feynman, in a New York Times interview.
The first step is usually the hardest.
connection.setRequestProperty ("User-agent", "GoogleBot/0.01");
I had never done this before and found it out by looking at the javadoc. I don't see how switching to Python would have made this easier.
Wikipedia's history of Java verions says 1.0, as does the request string in the article. [http://en.wikipedia.org/wiki/Java_version_history]
Was URLConnection available back then?
According to the URLConnection docs it's been around since JDK 1.0. [http://java.sun.com/j2se/1.3/docs/api/java/net/URLConnection...]
Did URLConnection.SetRequestProperty exist back in JDK 1.0?
The closest I could find were the docs for JDK 1.1.8 in a downloadable zip file, and yes SetRequestProperty existed back in JDK 1.1.8 at least.
Looking at the actual response, and the JDK 1.1.8 docs, he would probably have been using HTTPURLConnection (could not find HttpClient anywhere in the jdk1.1.8 docs) and even HTTPURLConnection in JDK 1.1.8 I could not find the string 'agent' anywhere on the page.
So yea, if the settings were there they were buried and not readily accessible in the documentation of the time.
setRequestProperty is there.
I know you saw the title and thought "ahahaha easy time to bash Java again", but it's really ignorant to do so.
"Most of Google is implemented in C or C++ for efficiency and can run in either Solaris or Linux."
and also
"Both the URLserver and the crawlers are implemented in Python."
So at some point they decided to try out Python for certain tasks.