This might explain why GigaBlast has a problem:
Because of bugs in the original Gigablast spidering code, the Findx crawler ended up on a blacklist in Project Honeypot as being “badly behaved” (fixed in our fork). That meant quite a bit of trouble for us because CDN providers, which are a very powerful hubs for internet traffic, put a lot of weight on this blacklist. Some of the most popular websites and services on the internet run through services like Cloudflare and other CDNs – so if you are in bad standing with them, suddenly a large part of the internet is not available, and we weren’t able index it.
extract from: https://web.archive.org/web/20190921180535/https://privacore...
Mojeek follows the robots.txt protocol so if a site doesn't want to be crawled by MojeekBot we respect that wish. There is also a generous crawl delay between pages on the same host.
Generally a 'badly behaved bot' will ignore robots.txt or hit a site too hard with requests.
Our bot uses a specific user agent which you can verify via DNS. https://www.mojeek.com/bot.html
What's the order of magnitude of this delay? milliseconds? hundreds of milliseconds? seconds? I'm curious what's considered 'polite' in this realm and how the various parties come to form opinions on this.
User-agent: bingbot
Allow : /
Crawl-delay: 10
https://en.wikipedia.org/wiki/Robots_exclusion_standard#Craw...Does this mean your spider is a fork of Gigablast? Is there some additional interesting technical information about how your code/infrastructure is set up?
https://blog.mojeek.com/2019/12/100-server-build-and-install...
We'll be writing about our tech stack in our next FAQs series; 3 of 4, this is 1 of 4:
https://blog.mojeek.com/2020/11/frequently-asked-questions-a...
I remember Gabe Weinberg mentioning on here that DDG became profitable pretty early on, so I'd be interested how easy it is to meet that threshold and how sustainable it is to create search competitors!
This means that they just need to sell enough advertising (which I think bing provides / forces them to use) so basically for DDG, more visitors = more revenue and breakeven means that the traffic pays for the costs they have.
If I were to create a privacy first search engine, I would likely not go general purpose and would instead try to focus on specific use cases first for specific audiences (that would be willing to pay). In practice this would look more like an extended version of DDG's instant answers than a Google-alike, but I think the world is convinced that general purpose web search is a) necessary and b) the only valid approach to finding answers to questions.
Many search engines already exists for specific use cases:
- www.bailii.org for UK law related searches like court verdicts
- tachiyomi.org for searching manga
Etc. Generally speaking, search engines that cater for a specific use case are simply called "aggregators".
You will always find:
><Search Engine> is <bad/good> and the results are way <worse/better> than <Another search engine>.
><Search Engine> is only a frontend for <Parent search engine> <but/and> does(n't) do their own indexing.
>Results for <Topic (usually technical)> (in <country>)? on <search engine> are garbage. Search results for <purpose> are poor.
What you will see as a response:
>You should use <search engine> instead! It uses <different parent search> instead so the results are better!
>I agree, I didn't like <search engine> so I had to go back to Google.
What you won't see:
>Nowadays I only use <search engine> for some search, most of the time I use <specific source(s)>.
>Maybe you should use specific searches for specific purposes?
The two that you cited are very valid, and I'm sure many people use specific sources for specific things (most doing it unknowingly via apps). The main thing seems to be that (on especially desktop) the single purpose all-encompassing web portal seems to be the way to find something. The phrase "Google it" is a catch all for 'do your research'.
DDG also uses Yandex for web search (mostly Russian content).
I've used DDG near-excusively since 2013.
Mojeek has been built from the ground-up meaning we do not depend on anyones technology. We have our own servers, crawler, index, ranking and so on.
So unlike most we are not using Bing, Google, Yandex or anyone else's. Our road to sustainability is a different one; we are a technology company.
Here are two posts on how we are funded and our business model; that post also covers privacy and (lack of) surveillance practices.
https://blog.mojeek.com/2020/10/who-funds-mojeek.html
https://blog.mojeek.com/2020/12/frequently-asked-questions-a...
Then, I'm wondering do you guys have the notion of "product brands" or something like that?
For example, when I search for "SaaSHub" - a product (I work on) that has been online since 2014 and have quite a few mentions around the web, the first result is some "random" Wordpress theme. i.e. if I'm searching for a particular product, I'd expect the product homepage to be the first result.
The feedback is appreciated and your example is one for a generic challenge that we are actively working on. I note that your site shows as link #2 so someone looking for your site/brand should see it; but still your point stands.
Do you disclose what goes into your ranking algorithm? I think having that transparency and perhaps being able to tweak some ranking parameters would be go a long way to being able to verify it's built in the best interests of the users.
A balance in ranking for those seeking information and those looking to promote content will always be an issue; at any level of transparency. Done right, more transparency can benefit seekers too.
This is a case of our bot blocking being overly aggressive, where bots tend to use quotes.
We will rectify this in the coming days, thanks again.
Given that you run your own spider, presumably you provided it one or more urls to get started spidering. I'm curious what those were.