I believe my algorithms are decent, but the biggest problem for Gigablast is now the index size. You do a search on Gigablast and say, well, why didn't it get this result that Google got. And that's because the index isn't big enough because I don't have the cash for the hardware. btw, I've been working on this engine for over 20 years and have coded probably 1-2M lines of code on it.
I wionder how much this is true, and how much (despite all our rhetoric to the contrary) it's because we have actually come to expect Google's modern proprietary page ranking, which counts more than just inbound links but all sorts of other signals (freshness, relevance to our previous queries, etc.).
We dislike the additional signals when it feels like Google is trying to second-guess our intentions, but we probably don't notice how well they work when they give us the result we expect in the first three links.
For me the experienced quality of Google search results gave have dropped massively since 2008, despite (and maybe even because of) all their new parameters.
When someone says this someone else usually immediately says it is because of web spam and black hat SEO.
But black hat SEO doesn't explain why verbatim doesn't work for many of us.
Black hat SEO doesn't explain why double quotes doesn't work.
Black hat SEO doesn't explain why there is no personal blacklists so all those who hate pintrest can blacklist them.
Black hat SEO probably also doesn't explain why I cannot find a unique strings in open source repos and instead get pages of not exactly webspam but answers to questions I didn't ask.
Back then Google was only going up against indexes and link-rings, not 2021 Google/Bing/DDG/etc.
You can see it with other search engines. I challenge you to come up with a Google query for which a first-page result won't be seen within the first 10 pages of Bing results for the same query.
(Bonus points if that result is relevant).
There's only so much tweaking that personalization and other heuristic can do.
But if something is missing from them index, that's it.
I was curious if you ever intend to implement OpenSearch API so that we could use it as default in browser or embed it in applications?
Also how can people contribute to help you maintain a larger index and/or keep the service going?
I'd pay 5-10$/mo for a search engine that didn't just funnel me into the revenue-extracting regions of the web like Google does.
[1]: https://michaelnielsen.org/ddi/how-to-crawl-a-quarter-billio...
Please elaborate. Is there a special relationship between Cloudflare and Google?
https://www.crunchbase.com/funding_round/cloudflare-series-d...
I'll see if I can track down the link but I remember somebody sharing a dump with me from Amazon that apparently was a recent scrape.
a) "Berlin":
1. The movie festival "Berlinale"
2. The Wikipedia entry about Berlin
3. Something about a venue "Little Berlin", but the link resolves to an online gaming site from Singapure
4. "Visit Berlin", the official tourism site of Berlin
5. The hash tag "#Berlin" on Twitter
6. "1011 Now" a local news site for Lincoln, Nebraska
7. "Freie Universität Berlin"
8. Some random "Berlin" videos on Youtube
9. The Berlin Declaration of the Open Access Initiative
10. Some random "Berlin" entries on IMDb
11. A "Berlin" Nightclub from Chicago
12. Some random "Berlin" books on Amazon
13. The town of Berlin, Maryland
14. Some random "Berlin" entries on Facebook
15. The BMW Berlin Marathon
b) "philosophy"
1. The Wikipedia entry about philosophy
2. "Skin Care, Fragrances, and Bath & Body Gifts" from philosophy.com
3. "Unconditional Love Shampoo, Bath & Shower Gel" from philosophy.com
4. Definition of Philosophy at Dictionary.com
5. The Stanford Encyclopedia of Philosophy
6. PhilPapers, an index and bibliography of philosophy
7. The University of Science and Philosophy, a rather insignificant institution that happens to use the domain philosophy.org
8. "What Can I Do With This Major?" section about philosophy
9. Pages on "philosophy" from "Psychology Today". I looked at the first and found it to be too short and eclectic to be useful.
10. The Department of philosophy of Tufts University
c) "history"
1. Some random pages from history.com
2. "Watch Full Episodes of Your Favorite Shows" from history.com
3. Some random pages from history.org
4. "Battle of Bunker Hill begins" from history.com
5. Some random "History" pages from bbc.co.uk
6. Some random pages from historyplace.com
7. The hash tag "#history" on Twitter
8. The Missouri Historical Society (mohistory.com)
9. Some random pages from History Channel
10. Some random pages from the U.S. Census Bureau (www.census.gov/history/)
d) "Caesar"
1. The Wikipedia entry about Caesar
2. Little Caesars Pizza
3. "CAESAR", a source for body measurement data. But the link is dead and resolves to SAE International, a professional association for engineering
4. The Caesar Stiftung, a neuroethology institute
5. Some random "Caesar" books on Amazon
6. Hotels and Casinos of a Caesars group
7. A very short bio of Julius Ceasar on livius.org
8. Texts on and from Caesar provided by a University of Chicago scholar
9. (Extremely short) articles related to Caesar from britannica.com
10. "Syria: Stories Behind Photos of Killed Detainees | Human Rights Watch". The photos were by an organization called the Caesar Files Group
So what I can see are some high ranked false positives that are somehow using the search term, but not in its basic meaning (a3, a11, b2, b3, d2, d3, d4, d6) or not even that (a6). Some results are ranking prominently although they are of minor importance for the (general) search term (a9, a13, b7, b8 -- perhaps a15 and d10). Then there are the links to the usual suspects such as Wikipedia, Twitter, Amazon, etc. (a2, a5, a8, a10, a12, a14, b7, c5, d1, d5); I understand that Wikipedia articles are featuring prominently, but for the others I would rather go directly to eg. Amazon when I am interested in finding a book (or use a search term like "Caesar amazon" or "Caesar books"). Well, and then there are the search results that are not completely off, but either contain almost no information, at least compared to the corresponding Wikipedia article and its summary (b4, b9, d7, d9), or that are too specific for the general search term (c1, c2, c3, c4, c6, c9, c10).That leaves me with the following more or less high quality results (outside of the Wikipedia pages): a1, a4, a7, b5, b6, b10, and d8. The a15 and d10 results I could tolerate if there had been more high quality results in front of them; but as a fourth and second, respectively, good result they seem to me to be too prominent. Also in the case of "Berlin" a4 should have been more prominent than a1, and a7 is somewhat arbitrary, because Humbolt University and the Technical University of Berlin are likewise important; what is completely missing is the official Website of the city of Berlin (English version at www.berlin.de/en/).
All in all, I would say that your ranking algorithm lacks semantic context. It seems the prominence of an entry is mainly determined by either just being from the big players like Twitter, Youtube, Amazon, Facebook, etc. or by the search term appearing in the domain name or the path of the resource, regardless of the quality of the content.
So my suggestion would be to lower the weight of the ranking of the domain, and promote sites which have a more recent update date.
Send me an email (contact in profile) if you want to follow up on this feedback!
Your general web search seems pretty good, although I've just given it a casual glance. I think your News search could be improved by just filtering the general search results for News-related content, since the "Ethiopia" content I get there is certainly Ethiopia-related.
In any case, an interesting product, I'll try to keep an eye on it.
Is this an issue of rate limiting, or request cadence? could you add randomness to the intervals in which you request the page?
Is it more complicated? do they use other signals to ascertain if you are a script or not like checking data from the browser (similar signals to the kind of things browser fingerprinting uses... e.g. screen res, user agent, cache availability, etc...) would it be possible for the browser to spoof this information?
I imagine rate limiting the IP address is the major issue... but could you not bounce the request through a proxy network? I've tried this with the TOR network before when writing web scrapers and had mixed success... seems like Google knows when a request is being made through Tor.
Perhaps you could use the users of your search engine as a proxy network through which to bounce the request for the scrape/indexing... This way the requests would look like they were coming from any of your users instead of one spiders ip address...Im not sure how cloudflare or any other reverse proxy could determine that thise requests were organic or not...
id be ok with contributing to a distributed search service so long as my cpu was not making requests to illegal content, and there were constraints put on the resource usage of my machine.
Sorry if this came off as all over the place, I do not know too much about the offense vs defense of scraping. These are just some thoughts...
Do you have a cost estimate? Also could you be more selective in indexing, e.g. by having users requests sites to be crawled.
It looks like the layout is hard-coded for a mobile browser, in portrait mode.
I tried looking up a game I'm interested in and the second results cluster from your search engine is a reddit thread about linux support for that game... I love this.
Great job!
I looked into it a long time ago and seem to remember there was a way to get access to registration records, but I imagine combining that with HTTP certificate transparency records would significantly increase your hostname list. Anything else?
Have you contributed your crawl data to common crawl?
Which is why the next big search engine should be distributed: https://yacy.net.
I know Google and Bing both use weird data-structure like BitFunnel
https://www.microsoft.com/en-us/research/publication/bitfunn...
I want to say - you don’t know what are talking about. But, it’ll be rude.
Hardware is much cheaper and powerful now compared to 2005.
AltaVista and Yahoo did that with browser plugins in the 90s.
I wrote a "meta" search utility for myself that can query multiple search engines from the command line.^1 It mixes the results into a simplified SERP ("metaSERP"), optimised for a text-only browser, with indicators to show which search engine each result came from. The key feature is that it allows for what I might call "continuation searches". Each metaSERPs contains timestamps in its page source to indicate when searches were executed, as well as preformatted HTTP paths. The next search can thus pick up where the previous one left off. Thus I can, if desired, build a maximum-sized metaSERP for each query.
The reason I wrote this is because search engines (not GigaBlast) funded by ads are increasingly trying to keep users on page one, where the "top ads" are, and they want to keep the number of results small. That's one change from 2005 and earlier. With AltaVista I used to dig deep into SERPs and there was a feeling of comprehensiveness; leave no stone unturned. Google has gradually ruined the ability to perform this type of searching with their now secretive and obviously biased behind-the-scenes ranking procedures.
Why is there no way to re-order results according to objective criteria, e.g., alphabetical; the user must accept the search engines' ordering, giving them the ability to "hide" results on pages the user will never view or simply not return them. That design is more favorable to advertising and less favorable to intellectual curiosity.
Each metaSERP, OTOH, is a file and is saved in a search directory for future reference; I will often go back to previous queries. I can later add more results to a metaSERP if desired. I actually like that GigaBlast's results are different than other search engines. The variety of results I get from different sources arguably improves the quality of the metaSERP. And, of course, metSERPs can be sorted according to objective criteria.
This is, AFAIK, a different way of searching. The "meta-search engines" of yesteryear did not do "continuations", probably because it was not necessary. Nor was there en expectation that user would want to save meta-searches to local files. Users were not trying to minimise their usage of a website, they were not trying to "un-google".
Today's world of web search is different, IMO. There seems to be a belief that the operator of a search engine can guess what a user is searching for, that a user who sends a query is only searching for one specific thing, and that the website has an ad to match with that query. At least, those are the only searches that really matter for advertising purposes. Serendipitous discovery while perusing results is not contemplated in the design. By serendipitous discovery I do not mean sending a random query, e.g., adding an "I'm feeling lucky" button, which to me always seemed like a bad joke.
The only downside so far is I ocassionally have to prune "one-off" searches that I do not want to save from the search directory. I am going to add an indicator at search time that a search is to be considered "ephemeral" and not meant to be saved. Periodically these ephemeral searches can then be pruned from the search directory automatically.
1. Of course this is not limited to web search engines. I also include various individual site search engines, e.g., Github.
did you ever think, let me just focus on Italy-relevant results? or job search only? or some slice of search.
The content quality will be higher and it's a lot cheaper.
Associating with that crank (responsible for recent freenode drama) is very off-putting.
What if instead of even trying to index the entire web, we moved one step back towards the curated directories of the early web? Give users a search engine and indexer that they control and host. Allow them to "follow" domains (or any partial URLs, like subreddits) that they trust.
Make it so that you can configure how many hops it is allowed to take from those trusted sources, similar to LinkedIn's levels of connections. If I'm hosting on my laptop, I might set it at 1 step removed, but if I've got an S3 bucket for my index I might go as far as 3 or 4 steps removed.
There are further optimizations that you could do, such as having your instance not index Wikipedia or Stack Overflow or whatever (instead using the built-in search and aggregating results).
I'm sure there are technical challenges I'm not thinking of, and this would absolutely be a tool that would best serve power users and programmers rather than average internet users. Such an engine wouldn't ever replace Google, but I'd think it would go a long way to making a better search engine for a single user's (or a certain subset of users') everyday web experience.
Google's not exactly working against the echo chamber problem, and I think that's because to do so would be to work against its own reason for existing. There are two goals here that are fundamentally at odds with each other:
1) Finding what you're looking for.
2) Finding a new perspective on something.
A search engine's job is to address the first challenge: finding something that the user is looking for. The search engine might end up serving both needs if they're looking for a new perspective on something, but if these two goals ever come into conflict with each other the search engine does (and I would argue it should) choose the first goal. Failing to do so will just lead to people ignoring the results.
If you're indexing forums or social media, the same site is going to give back the bubbled responses, possibly without the person even being aware they're in a bubble.
https://www.google.com/search?q=%22BATF%22+guns&client=safar...
https://www.google.com/search?q=%22ATF%22+guns&client=safari...
Interestingly, back then, Google was big on neutrality and refused to do anything, stating that it reflected the way people used the word. It was finally addressed using "Google bombing" techniques. Something that Google didn't care much about back them because of its low impact.
Similarly everyone maintaining your their own index is cumbersome overkill in redundancy, processing power, and human effort in return for a stunted network graph which is worse for all metrics people usually actually care about. In terms of catching on even "antipattern search engines" that attempt to create an ideological echo chamber would probably catch on better.
Short of search engine experiments/start up attempts the only other useful application I can see is "rude web-spidering" which deliberately disrespects all requests to not index pages left publicly accessible as search engines generally try not to be good tools for cracker wardriving for PR and liability reasons. It would be a good whitehat or greyhat tool as doors secured by politeness only are not secure.
Capital is the huge barrier to entry today:
Larry Page's genius was to extend google's tech, consumer-habit and PR barriers-to-entry into a capital-based advantage: massive geo server farms, giving faster responses. Consumers have a demonstrated huge preference for faster response.
Basically, SEO. SEO is the real problem, not search engine algorithms. Those algorithms are a result of the arms race between Google and black-hat SEO BS. Remove SEO and search engines work just fine.
While that may be good for most people, there is still a lot of power and utility in simple keyword-driven searches. Sadly, it seems like every major search engine has to follow Google's lead.
Like, instead of trying
PDP11 emulator
PDP-11 emulator
"PDP 11" emulator
PDP11 emulators
PDP-11 emulators
"PDP 11" emulators
PDP11 emulation
PDP-11 emulation
"PDP 11" emulation
Basic NLP can do that a lot faster without introducing a lot of problems.I do think Google currently goes way overboard with the NLP. It often feels like the query parser is an adversary you need to outsmart to get to the good results, rather than something that's actually helpful. That's not a great vibe. However, I think the big problem isn't what they are doing, but how little control you have over the process.
Early Google was a breath of fresh air compared to the stemming that its competitors at the time did, but nowadays even putting search terms in quotes doesn't seem to return the same quality of results for these types of queries that Google used to have.
Disclaimer: I work on Google search.
One porridge is too hot and the other is too cold. I know Google could find a happy compromise here if it wanted to. In fact, I bet there's some internal-only hacked-together version that works this way and actually gives an acceptable user experience for the kind of people who have shown up to this thread to show their dissatisfaction.
Two results not containing "eggz" at all. Two results containing "eggzackly<punctuation>this" Two results containing "eggzackly" but missing "this".
Google Search is broken. It no longer does what it's directed, it just takes a guess. I suspect part of this is because someone decided that "no results found" was the worst possible result a search engine could give.
You can search for, say, "cycling (insert product category here)" and get motorcycle related results. Why? Because to google "cycling" = "biking" and "motorcycles" are "bikes", bob's your uncle, now you're getting hits for motorcycle products.
Every time I try to do a very specific search I can see from the search results how google tries to "help", especially if the topic is esoteric. The pages actually about the esoteric thing I'm searching for get drowned in a sea of SEO'd bullshit about a word/topic that is 1-2 degrees of separation from each other in a thesaurus. I'm sure someone at google is very, very proud of this because it increases their measure for search user satisfaction X percent.
It does this thesaurus crap even with words in quotes, which is especially infuriating.
This is called "stemming" and is not sensibly approached with machine learning.
It originates from the disease 'the pox'.
No-tracking and independent from the start. Now at 4.6 billion pages with own infrastructure and IP. Went to market in 2020 with contextual ads and API. Self-disclosure: CEO
Case 1: You just want the name of the website, or an article, example "Facebook" -> fb.com, "Gordan Ramsay" -> Wiki/official website/Celeb gossip website you are good. Not much competition here.
Case 2: You are looking for something technical like "GNU rnano CVE-abcde"/"OpenBSD ARM64 Qualcomm Wifi driver not working", you are again in the fine territory, not much if any money to be made here so very less competition. There will be the official forums, websites, maybe some conference websites in this category.
Case 3: "Chicken potpie recipe", "How to be more organised": This is the category where people are trying to game the SEO algos. How the hell do recipe websites with 27 popups, 12000 word essay on the secret family history ends up on top ? There are a huge number of passionately made simple recipe websites but they have to be "found" by us. For the second query I mention about being more organized I think most people are looking for some sort of a review article which looks at some various schools of thoughts regarding discipline, cleanliness pointing to further resources and exploring the why and what to do for this. Here the search engine needs to determine the context of the query which is fairly abstract and then the internal heuristics it uses are supposed to drive it to a meaningful list of websites. Maybe the average joe would like to click cosmopolitan's article but I would never do that. Based on my previous click history maybe google should determine what I kind of links am I looking for. But when they figure that out they'd much faster use this behavioral insight for advertisers. A great search engine is basically a primitive personal librarian, I'd pay a yearly subscription for one.
The internet is vast and it has stuff that I don't know about. How my 7 word abstract query is gonna get me there is the question mark. Also, for a lot of queries the top results can be plagued by spammy/fraud results which are on top because they managed to trick the SEO algos. These bad actors were not as prevalent for 2005 google.
It's a bit like the car industry - you could run a startup from your garage in the early days but you need titanic amounts of capital to compete now thanks to vertical integration.
Major governments and billionaires can compete but everybody else is locked out of the market (most "startups" use bings index).
Running a search engine in your garage is feasible today because hardware and connectivity have improved much faster than the size of the WWW.
Also, that user data is used to improve search results and mitigate webspam that didnt exist in 2005.
I’ve thought the same about pre-ad Twitter and Facebook.
Early on, startups with free services look a lot like non-profits and just maximize user benefit to grow. The problem is they’re not non-profits, and have to make money at some point. That has tended to mean ads.
I’d easily pay, say, $9/mo to have access to an ad-free search engine that made me feel the way 1999 Google did.
But even charging users $21.33/mo for an ad-free search experience most likely wouldn't be enough. By providing such an option, you'd greatly reduce the value of the remaining Ads pool.
The optimistic perspective on this is that if you are one of the users with disposable income, you're essentially subsidizing a great search engine and a suite of other tools for the less well-off ones.
I’d bet there’s some way to characterize what I and others liked about the earlier web and create a search engine that just worries about that stuff. I’d pay $9/mo for whatever 1/3 of Google’s spend per user would get me. That’s not to say this thing would “beat” Google, but it could profitably exist.
And then I'd guess the 20 remaining users will still complain because 1999 Google is a nostalgic memory impossible to recreate without a 1999 internet for a 1999 self to live in and has little to do with raw search quality.
If you took an exact copy of Google circa-2005 and had it crawl today's web, you'd probably get mostly "SEO optimized" irrelevant blogspam.
The early Google (and other even earlier search engines) were invented for an Internet world which, if not pristine and pure, was at least mostly fairly legit content. Today's Internet is probably 90% deliberate spammers and scammers.
Part of the problem is that there's a lot more low-quality content to wade through now than there was in 2005. I think the Google of 2005 would have trouble delivering quality results today also.
I wish there was an easy way to filter ALL search results, by permanently excluding specific websites, and/or keywords.
Surely there has to be some browser extension that does this...
Not got round to trying it yet though.
I don't use Google anymore to search unless I really need to. The algos they use today are not the same classic ones that actually returned results.
(Plus I can directly go to the wiki page by using "!w", "!gm" for google maps, etc.)
What would you attribute to their modern 2021 success then? Just throwing a ton of money at amazing engineers to hone in their complex algorithm to tweak it to still return what us humans quantify as "good" results? Especially if they are waning through a sea of low-quality content as you say.
Edit: Forgot to say that I work on Brave Search.
Explanation is also found here with screenshots: https://search.brave.com/help/independence
I get 84% personal (browser-based), 87% global (which means we hit Bing only 13% of the time from our server side).
Brave search has been my daily driver and it works wonderful.
In my experience it is better than Google at what it does if I'm looking for long-form texts (exception being scientific/peer-reviewed articles, Google tends to shoot me those for the type of queries I make on Marginalia), but is very much complementary rather than a replacement.
Right now the biggest problem with Marginalia is that it has a fairly uneven quality level. For some queries it's absolutely incredible. For others, it doesn't really provide much useful results at all. I do think it's possible to even that out a considerable bit, to make it more viable for general queries. It's never going to be able to answer every query, but it probably could answer a lot more than it does.
Disclaimer: I study for malicious JS stuff.
And that specifically blocked Pinterest, Quora, most non-personal “blogs”, etc.
People suggest DDG ! operators, but I don’t want to use a site’s (bad, single-site) search box. I want a multi-site SERP that only displays results from known good sites, which are customizable.
I've done similar optimizations elsewhere to counter Google's trash results, e.g., I've been beefing up my personal recipe database, with the goal being that I can avoid a google search altogether whenever possible, only hitting google as a last resort.
More and more I wonder, with the modern internet, is it even a feature that the whole web is indexed? Might be a bug.
Search engine as a social media platform? If I follow you, now I can search in your indices?
Also gmail, used to have the best spam filters out there, now it's utter crap. Emails from my google analytics account, for whatever reason and disregarding how many times I have clicked on "Not Spam", go to spam, and it's their own service; while messages who are textbook spam ("Hi, I just got some inheritance ...") go to my inbox.
AI (in its current state) is crap, when is the industry going to accept these are the emperor's new clothes.
Thus, the Google of today, which is optimized to extract that money from us.
An easy way to become way better than google — detect google ads on pages, and penalize these pages in the index. For obvious reason, google search is incapable of doing so.
“Here, you see, it takes all the running you can do to keep in the same place.”
-Lewis Carroll's Through the Looking Glass
With content copying, shuffling and AI generating, I am afraid we are on the cusp of auto content generators passing some restricted Turing test where readers really think it's an actual human that wrote it.
As for me, I leant that for certain "hot topics", simply doing a generic search on Google is not a good idea anymore.
Then, check Google's ranking of the page. If it is much higher than it seems the page should be, assume the page is being SEO hyper-optimized and penalize the page proportionately.
Basically, using the variance between Google's model and your model as an indicator of an SEO spam page.
[0]: https://brave.com/privacy/browser/#web-discovery-project
We also need to be aware that when we remember past times it usually carries a romantic, nostalgic note. Web is very different than it was 15 years ago and the problem of search has evolved.
What you are looking for is basically 'grep for the web' but it is just one facet of search that we use today. 15 years ago you would not get an instant answer to a question like you do today and many users would not be able to live without that today. There are also maps and location based answers, all sorts of widgets like translation etc. Also world became more polarized so an objective best search result became more difficult to produce, specially for events covered in news, which means bias inevitably starts to creep in.
This is not to say that Google is good or bad today, it is what it is and they are doing best they can. Startups like ours see an opportunity on the market, in large part to help savvy users find what they want.
[1] https://kagi.com
“Information Neutrality is the principle to treat all information provided (by a service) equally. The information provided, after being processed by an information-neutral service, is the same for every user requesting it, independent of the user’s attributes, including, e.g., origin, history or personal preferences and independent of the financial or influential interest of the service provider, as well as independent of the timeliness of information."
I wrote about this in relation to search [0]. We need to be allowed more freedom to choose search engines and services. One (default or selected) choice for search is unhealthy. We shouldn't have to choose between Google or Bing; DuckDuckGo or Startpage; Brave or Ecosia; Mojeek or Gigablast ..... Personally I use all 8 of these and more, as also explained [0].
[0] https://blog.mojeek.com/2021/09/multiple-choice-in-search.ht...
I like Firefox's UI when searching, where you can select the search engine of choice while typing a query.
I like customizable metasearch engines like searx, I think it is a phenomenal idea. I wish more niche engines would implement OpenSearch so that they could easily be added.
I have considered just making a simple web page with search boxes for multiple engines for personal use as a default home page, but there's friction and again, lots of engines I'd like to use don't implement OpenSearch.
I wonder if there's some novel UX approaches to this out there. Meta search engines seem to be the best way so far to do it but there's the problem of customizing ranking, relevance of results and the like that just compounds the problems users experience.
https://www.burda.com/en/news/cliqz-closes-areas-browser-and...
https://news.ycombinator.com/item?id=23031520
https://0x65.dev/blog/2019-12-06/building-a-search-engine-fr...
Cliqz is now Brave Search, I use it for all my devices, it's great.
Works better than DDG and sometimes better than Google.
I only hash bang every 100 searches or so, most of the time Google doesn't have it either. It's just to make sure.
Also user preferences have changed in the last decade or so. I know millenaials and users in their late 30's or early 40's still yearn for the old web where they would type a search term and correct results would astonish them. However, younger users tend to gravitate to videos and that is why a large portion of the google results are now video results.
It is called Poe's law, and Google returned it at #4. Bing or Duckduckgo don't have a clue...
2) They have a years of user's data, like for specific term, they see what users clicked most, so they see which results were perceived as most relevant. It is hard to catch up if you dont have such data.
3) They developed anti-spamming tools during the years of fighting against SEO-spammers.
My problem there is that I don't expect or want my search engine to do that. The counter case is where I remember a quote from and article and want to find the article. Old Google would help me find matching text and I could quickly find the original article. Current Google will try to interpret the text and give me some nonsense based on that.
AI has ruined other Google features... the "search by image" feature now analyzes the image, returns a generic tag like "woman", and shows me the wikipedia article on women as the first result.
Old search by image had tineye like functionality and you could find the source of images.
> It is called Poe's law, and Google returned it at #4. Bing or Duckduckgo don't have a clue...
Interesting, I was looking for a good benchmark like this. For me Google returned it at #5 with an image/related terms carousel before it which places it physically more around #7 on the page. Brave Search (never tried it before today) puts Poe's Law at #8. So Google is still better.
But the other results are mostly worse (IMO) on Google. Here are the first 8 results:
- 175 Bad Jokes That You Can't Help But Laugh At - Reader's (rd.com)
- 57 Hilarious, Silly Jokes No One Is Too Old to Laugh At (bestlifeonline.com)
- 145 Best Dad Jokes That Will Have the Whole Family Laughing (countryliving.com)
- Sarcasm, Self-Deprecation, and Inside Jokes: A User's Guide (hbr.org)
- Poe's law - Wikipedia (wikipedia.org)
- Managing Conflict with Humor - HelpGuide.org (helpguide.org)
- 175 Bad Jokes That Are So Cringeworthy, You Can't ... - Parade (parade.com)
- Encouraging Your Child's Sense of Humor (for Parents) - Kids ... (kidshealth.org)
And here are the first 8 results from Brave Search:
- phrase requests - Is there a word for "pretending to joke when ... (english.stackexchange.com)
- Joke - Wikipedia (wikipedia.org)
- “Are you joking or serious?” – The Caffeinated Autistic (thecaffeinatedautistic.wordpress.com)
- How do I tell when people are joking or being serious? (reddit.com/r/socialskills)
- be a joke | meaning of be a joke in Longman Dictionary of (ldoceonline.com)
- Quote by Ricky Gervais: “If you can't joke about the most (goodreads.com)
- How can you tell if someone is joking with you or not? (quora.com)
- Poe's law - Wikipedia (wikipedia.org)
-----
edit: I did not count to 8 correctly the first time. Fixed that.
Search engine isn’t singular, it’s plural.
(1) Search engine for something I know exists.
(2) Search engine for finding something new.
There’s a market for both, but you don’t have to solve both problems with the same product.
Sometimes I switch to Google for the former, but the latter works well enough for me that I don’t care what else Google would’ve shown me.
More often than not, my feeling is Google would only have shown me more ads in addition to whatever I could already find elsewhere.
instead of discussion forums and Q&A sites, everyone's on facebook/twitter/discord/slack/snapchat/tiktok/etc... none of that is really very google friendly
online marketing and SEO is a much larger industry now, so with less (by % of total) searchable content generated by people (which is on social media) a lot of the high-ranking content that appears in search is highly optimized marketing
then you have other kind of weird things like... half of all internet traffic being bots
Now that Google exists, you can't create another one. There's only room for one.
Another thing is the rise of "content sites", like this one (Hacker News). I'm sure YCombinator doesn't like getting hit by dozens of crawlers. The impulse to ban everything that crawls except (Google|Bing|Baidu|VK) is too great.
A lot of alternative suggestions are being thrown into this discussion. Let me throw in mine: Reverse the concept of the "crawler". Instead of following links around the internet randomly, require sites to register with you and request to be crawled and/or submit a sitemap. It would be hard to get started, but once something like this gained momentum, I believe that there's room for several of these reverse-search-engines to compete.
Just let me type stuff into the search box -- including typo corrections and modifications to what I'm searching for -- and hit ENTER to start the actual search.
When I'm ready to start my search I'll hit the fucking ENTER key. Stop annoying me with your stupid assumptions about what I'm looking for.
This ONE THING is why I switched to Webcrawler.com two years ago. I type in five or ten words with ZERO craptastic guesses flashing around on my screen, hit ENTER, and THEN it returns what I'm looking for.
DDG also doesn't support showing a site's basic structure in the search results (ie, the card of a company's website with Products, Contact Us, Support, etc) and the preview text is garbo as well...it reminds me of 1990's era electronic card catalog search excerpts.
I look at the first page or two, give up, search google. While I have to hunt a bit in the results, I do eventually get what I wanted.
Even in 2021, despite how bad it's become, it's still miles ahead of other competitors.
And another private company is not the answer I believe. We need something more drastic, an open-source search engine organized as a genuine non-profit organization. Something like that. Otherwise, whatever replaces Google will just turn into another Google as soon as it gets any momentum.
In that era, Google would return a match based on words that appear in the links to a URL but not in the article itself, meaning that it was easy to produce "Googlebombs". For example, from 2005-2007 the top hit for "miserable failure" was the Wikipedia article for George W. Bush.
See https://www.screamingfrog.co.uk/google-bombs/ for some of the "better" ones.
I heard HN constantly crying over its deteriorating quality, but I am not noticing it that much, not better not worse, it just does its job.
To create 05 Google, it is easily billions of dollars and years of investment, before people will treat you seriously.
The reason we didn't get 05 Google could only because it is not profitable. Some nation state attempt to demonopolize the search engine business might work, but I didn't expect any for profit organization to easily attempt doing this, let alone individual hobbyists
The parent company, Tiscali, was a huge hit in the 1990s, as it provided internet access to millions of Italians. It went through some struggle for several years, but lately the original founder, Renato Soru, came back to run the company.
The company is based in Cagliari, the capital of Sardinia, Italy.
Sure, it feels great when the engine guesses something like that correctly -- but it comes out worse overall for the plentiful cases where you have to try to compensate for it guessing wrong.
I can only think of examples where I want personalization. What's an example query where it interferes?
If i wanted something more relevant to me, then i would specify what aspect of relevance (country, gender, age etc...) i would like instead of playing the guessing game.
If that's true, then I don't think you are a typical search engine user.
The personalization should just be used for defaults. You can always make a more specific query to focus on aspects you are interested in.
I wouldn't use it for everything but sometimes that is the exact behavior that I want. I'd use duck duck go for more general searches.
Imo, 2005 google got initial traction because of its tech forum post indexing, as I remember my switch to it was because it became an extension and then replacement for manpages. In that sense, what made it good was it reflected the consensus of what its incredibly influential userbase thought was important and just managed that really well. The demographic impact of the U.S. Gen X all using it at once didn't hurt either.
The equivalent today, as a lot of us say, is that blockchains are in the 1997 internet phase, and the service that makes the content of those as navigable as the 90's internet, will likely grow in a similar way.
Search that provides young people with privacy and freedom to pursue their true interests will be the dominant strategy. Its success will be because it's a product that rides growth, and not because it "solved a problem." Imo, we all index too much on the privacy pattern because the freedom pattern is too risky.
What's changed since that time are the maturity of things like Bloom and other probabilistic filters, Apple's private set intersection, differential privacy, zksnarks, and everybody you'd ask an opinion from now gets their content through mobile devices. Apple's ecosystem is equipped to do this kind of search, but they're too exposed politically to get into it. Meta will likely go there, but nobody's going to trust them willingly.
A protocol that generated a cryptograpically strong anonymous index from your browsing - and instead of putting it on google's servers, it was on a chain, or the content index information and its evolving consensus score was included in something like a DNS record - may still unseat these ensconced interests. IPFS and other P2P or torrents might do something like that as well. Blockchains maybe good for that consensus/desire score.
It's not something you architect and design top down that has to solve all cases, it will be just another useful product that grows while riding a demographic change. It would be on the level of inventing HTML/HTTP again, which, when you think about it, was just another dude making a thing he needed.
Rather than being told "No, there are only eight pages of results on anything in the goddamned world. Really. Would I lie to you?"
Gigablast Search Engine - https://news.ycombinator.com/item?id=29421898 - Dec 2021 (10 comments)
* Don't use JS * Don't use Google analytics * Don't weigh more than a few kB per page * Don't show any sites with ads
That would be a place to begin.
Because the universe being searched isn't the internet of 2005 and earlier, and because user expectations have moved on, too.
Plus the index expense.
For example if my search term appears in the URL I can almost guarantee I don’t want that page.
I'd gladly pool in some of my CPU time if it helps build a better search.
Well, it's wikipedia. So just create a search engine for that, since their search sucks rocks.
Knuth's "Searching and Sorting" volume desperately needs an update.
https://www.amazon.com/Managing-Gigabytes-Compressing-Multim...
https://www.amazon.com/Information-Retrieval-Implementing-Ev...
https://www.amazon.com/Introduction-Information-Retrieval-Ch...
Ask HN: Has Google search become quantitatively worse?
https://news.ycombinator.com/item?id=29392702
Inviting all the paranoid/speculative/hearsay/personal experience responses. Lame Ask HNs!!!!!
DDG is pretty useless though unfortunately.