Gigablast Search Engine, Now Open Source (C/C++)
gigablast.com
gigablast.com
Some facts about the engine:
The code compiles into a single executable file which can scale on thousands servers.
It is easily configurable and has a nice documentation [2].
The code is very stable, it works in production since 2002.
Document processing is done using plugins, so you can write a plugin for any type of documents.
---
I would like to see a search engine based on this in the dark-nets, particularly in I2P.
[0] - http://www.prnewswire.com/news-releases/gigablast-now-an-ope...
[1] - https://github.com/gigablast/open-source-search-engine
I really really needed that just right now! You helped me soo much =) Thank you sir!
EDIT: Gotta say, this has some very useful pieces of code. I'm working on a niche-specific crawler and am battling the url stripping/cleanup part of it. This is very useful: https://github.com/gigablast/open-source-search-engine/blob/...
Unfortunately the document filter in questioning dose spawn child processes, so the normal way of using fork() and a monitoring process was not working. However using ulimit like this should work: https://github.com/gigablast/open-source-search-engine/blob/... . Hadn’t thought about spanning a new shell and let it have control like that :)
Old habits perhaps? When I look back at it I remember that my first books on C were full of problematic sprintf and strcpy use. It may then easy to continue using what you first learned, even when you know better. It basically the "Baby duck syndrome"[0] for C functions.
0: http://en.wikipedia.org/wiki/Imprinting_(psychology)#Baby_du...
Interesting read, its history: http://www.gigablast.com/press.html
The great thing about this project is that it comes with good documentation for administrators and developers who want to extend it. As Gigablast has been sold to enterprise customers.
Admin Docu - how to build the source, troubleshooting, etc.: http://www.gigablast.com/admin.html
Developer Docu - even explains how to use Bash, GIT what to do on hardware failures, etc.: http://www.gigablast.com/developer.html
Two Search Engine features are currently disabled because of code overhaul: Boolean query support & Spellchecker. As Google is removing more and more such advanced features from its search engine - "+" anyone. It would be great if these features would celebrate a comeback, either from its original developer or with the help from the open source community.
Thanks for open-sourcing it.
BTW, long ago I hoped Gigablast would become a popular google competitor; no such luck. I remember asking Matt if I could provide an official IE toolbar (when they were the rage) he declined; sadly. My hope has shifted to duckduckgo.
I look forward to forking!
const char *CountryCode::getAbbr(int index) { if(index < 0 || index > s_numCountryCodes) index = 0; return(s_countryCode[index]); }
https://github.com/gigablast/open-source-search-engine/blob/...
Lots of details about its development on WebMasterWorld, it only uses a handful of servers.
I couldn't get it to compile on my ubuntu 13 machine with out some errors and warnings, so I forked it and made some changes. i don't know git very well so i don't know how to merge, etc.
However, given the scale of the project and the fact that the code has been in production for more than 10 years, it's more likely the errors you faced were due to:
- your local environment not being configured ideally, or
- "configuration code" that you did not modify. :)
It says in html/admin.html to just type make to compile.
You will need the following packages installed
apt-get install make
apt-get install g++
apt-get install libssl-dev (for the includes, 32-bit libs are here)
1. Run 'make' to compile. (e.g. use 'make -j 4' to compile on four cores)I am yet to try installing libssl-dev though as I don't have root access on the machine I was testing on.
The page is no more, Archive.org has no copy (due robots.txt flag) but Google has still a cached copy of the blog:
http://webcache.googleusercontent.com/search?q=cache:9lS6Ngk...