A search engine that removes the 1 million most popular web sites from its index
millionshort.com
millionshort.com
I actually find myself narrowing my searches by sites I know will have reliable information, like this one, reddit, certain forums based on the search topic, etc. I think there would be some real value in creating a search engine that was very selective about the sites it crawls. Honestly, crawling the comments from the best user-participation sites on the web (reddit, HN, SO, quora, etc.) would probably make for a very useful search engine.
#1, #2 and #3 results are spam pages on a legitmate site.
Since I remembered Blekko (met Greg, the founder guy, in SF once) I pulled it up and we searched for yunnan, a Provice in China, as a test. Most of the results were for commercial tour operators or thinly veiled redirects for such.
The million-missing one on the other hand turned up more interesting or 'bespoke' content.
I am ignorant of such matters but would have thought a spamassassin-style Bayesian model based on sentiment analysis, advertising frequency, update frequency, hosting location, content originality or any similar clump of readily obtainable metrics would be enough to usefully cull the vast majority of the useless modern stuff.
I mean, if I want commercial tour operators, I'll tell the search engine by typing something like "prices" or "companies" or "costs" or whatnot.
The other realization we had, doing this on an iPad, was that search engine interfaces positively suck. They're still stuck in the 90s. With a touch-based interface, there should be a more interactive model for query refinement than text editing. Like, uncheck [x] commercial sites.
For example, sites like yellowpages.com, whitepages.com, superpages.com, zillow.com, citysquares.com will all pollute basic searches that look like job descriptions ("custom cabinetry new york", for example).
(I complained to DDG about this a while back, and it looks like they have added some negative boosts to some spam sites.)
I searched for "180sx" on this site and the first ENTIRE PAGE of links were amazingly relevant and from sites I'd never seen before. Seriously useful information -local Australian body part suppliers, build logs.
I google car-related stuff all day and this is the most useful stuff I've seen in ages. From the first page of results.
Your idea started out good. Sometimes I want to research a product before I go looking for a merchant. Blocking all sites that contain ads? Just use adblock if you're that fanatical.
I think it's a good idea. It'd certainly help with the problem of blogspam etc.
Even of you refined your criteria, what do you have against sites that try to recoup costs of hosting or content? Or, are you referring to completely different sites, lumping them all under "commercial sites" label?
More generally, I don't have anything against any sites, I'm trying to come up with a way of viewing the internet that would have a high signal to noise ratio. It seems plausible to me that excluding all sites with ads would improve the signal to noise ratio over what we have today. If you think the signal to noise ratio would be improved further by carefully including some sites with ads (SO is a good example), can you take a stab at defining the criteria?
Maybe include only sites whose ads are deemed acceptable by adblock? But then we might include many content farms...
As the owner of a search engine I can tell you search ain't cheap.
Bikeshedding I say!
If we're just making something interesting and different, not shooting to replace all search, imperfect definitions are a wonderful place to start.
A mix of affiliate-driven content sites, abandoned blogs & spam sites, based on some of the searches I did. The broader the category, the more likely I think you are to stumble on something relevant and helpful.
(Also, .gov sites don't seem to have been removed from the index. A search for "type 2 diabetes" still brings a number of results from NIH.gov, mimicking what Google serves up.)
Tiny complaint: Its default option is "Don't remove any sites". Kind of misses the point imho...
Other complaint: If country is set to e.g. Switzerland, you see only .de and .ch domains in the results! Is there no "worldwide" setting?
Which is a shame because I really enjoyed searching with it.
A How To should rank better for "Make a cake" than a sales page or a review page.
A Review should rank better for "Best SUV" than a Table Of Contents page.
Just because you are the Underdog doesn't mean you should win. You should have a fair fight, but a good result is a good result.
I hate much of the stuff in Wikipedia. But there are some pages that were amazingly well written.
I hate eHow, but there was a brief time when they had experts in the fields they were writing about writing really great content. Those posts should do well.
Later we may let you turn off Right or Left Leaning articles. We may expose the feature we have that returns only easy to read results, a feature designed for ESL, and Youth searches.
But we will never release a blanket no more top million sites.
Neither page popularity or query popularity are necessarily proportional to domain popularity (eg, *.github.com). Ruling out the most popular domains is therefore, I suspect, neither good or bad in terms of the quality of results it produces on the whole. Sometimes it will produce better results, sometimes it will produce worse, sometimes the same.
If a search engine/tool is going to add value, imho, the very difficult problem that it must solve is to improve the quality of results. Unless I'm missing something, I don't see that here (yet).
(I like big robots. I searched for "mech." #9 was a mech-based browser MMO I'd looked at briefly a couple weeks ago.)
https://millionshort.com/search.php?q=Adam+West&remove=1000k
www.adamwest.com isn't removed because it's ranked 3,400,000th or so, using alexa as a reference (http://www.alexa.com/siteinfo/adamwest.com)
Some of them may rank them in ways more useful to more people of course. Getting ranking working well is pretty much the challenge in making a good web search site, and is how Google originally vaulted over it's competitors.
Asking to know "how the data is collected, collated, and sorted" is… asking for a lot. For one thing, it's pretty much 'trade secrets' -- lots of people would really like to know the answer to those questions about Google, but it would probably take a multi-volume book to answer it completely, and Google generally prefers not to share the details, as they are what makes Google what it is (and would also make it easier for people trying to spam google results)
The formula changes based on the type of search, and the type of results we found, but the indicators are all exposed.
Also, I'm not asking for every last detail. Even Google has shared the broad outline of their PageRank method. I don't think it's at all unreasonable to expect some explanation of the basic approaches involved.
And I would guess this would be closer to vanilla page-rank than Google's secret sauce...
I don't know who's behind that search engine, but thank you!
I ask because my two top results for 'bitcoin' are bitcoinbrasil.com.br and bitcoin-today.yoyafi.com