Yes, it needs to get better. But it's a good start, I think.
Yes, it needs to get better. But it's a good start, I think.
Keep in mind that Google is not a propaganda machine, it really really _really_ solves problems and finds things. Any search engine willing to compete will have to do this.
I'm not saying it is impossible, but that it is not sufficient to be "advertising-free" to be good.
What are some of your plans for search engine improvement? Ad-free aside, why should I fund your project? DuckDuckGo seems to offer itself up as a great alternative to Google and I have enjoyed its clean UI immensely for several months now. What is my incentive to help drive your project?
This is currently written in Pascal. I'm in the process of rewriting it JavaScript/Node.js. First part for the rewrite is the ranking. That should be done in about a week or two.
You should write a blog post about how you write a search engine in Pascal. That would at least reach top spot on HN!
But I know and follow the YaCy-community. Like myself the lead person of YaCy is located in Germany. We even share the same first name. :)
JavaScript seems to have the biggest community at the moment. Plus I kinda like Node.js. That's why I'm going that way.
I encourage you to continue in your efforts, and I will add some suggestions on how you might proceed which will get you further along.
First, in order to scale, it will have to be able to run on 'n' servers. So look at ways of breaking up your index across multiple machines and operating on the index in parallel.
Second, crawling can often be a bigger challenge than providing search results, so work on building a system that can crawl a bunch of different web sites. Folks like Octopart have shown that topical search engines are really useful. So consider putting together a billion document index of say "news sources".
In the 'everything that is old is new again' theme, consider building simply a 'blog search engine' which is one that focuses on various blog content. The old Technorati did that and later was subsumed and its become harder for blog writers to rank in Google's results, so perhaps there is an opening for a place to search about who is talking about 'x' for some X.
I happen to know you can run a pretty decent crawler and indexer for about $2M a year :-) so consider targeting your donation rate to hit about $150,000 a month.
I'm already considering making "special" indexes for blogs, news etc. like you mentioned. In fact that same suggestion came up in a conversation I had with someone today.
The crawler/indexer is pretty fast as it is. I could crawl about 2 billion pages/month for a cost of about 800€ / 900US$ per month. That's on a 3.4GHz quad-core machine with a 1gbit/s connection and 32gb RAM. I tested that once. Works fine.
I can only imagine that your parser does a lot more than mine does. Crawling is CPU-bound for me. So the more processing the parser has to do, the slower the crawling would be.
Much of what we crawl is shared with the common crawl folks so that can help you with your choosing perhaps. Also that corpus provides a good way to test indexing and ranking algorithms. Things I like to test are queries like "Who are the Cardinals?" which can catch both sports references (Arizona and St Louis Cardinals) and church references (Pope Francis just named/blessed/appointed a bunch of new Cardinals). Often times news sources or blogs can help identify which interpretation of an ambiguous word is probably more relevant at the moment.
Interpreting ambiguous words for ranking is beyond what I can do at the moment.
I saw some references to round-robin in your code, I suppose you use it to hit a domain only every x milliseconds during a crawl.
Maybe additionally use Open Directory Project as start list: http://dmoz.org