Two upstart search engines are teaming up to take on Google
wired.com
wired.com
I just don't see how an LLM fits my search habits. I tried using ChatGPT for search purposes, and it was dreadful. Incorrect hours for businesses, gibberish results for the copy-pasted error messages, it was an all-around failure.
Have you people been typing complete sentences into Google all these years? Is that where this whole "LLMs will replace search" thing is coming from?
I considered Kagi. I like the concept, and I'm willing to pay for search, but Kagi costs more than I'm willing to pay.
- An autogenerated info card that usually is regurgitating wikipedia.
- Several SEO blog spam posts that have the exact query as the title and kinda answer it while trying to sell you something. No sources for more info of course, just a link to the store or whatever.
With this landscape, LLMs for search kinda make sense. In a way we are already 90% of the way there.
LLMs are indeed not the right tool for finding opening hours or phone numbers (at least for now), but I actually find them more useful than Google for error messages. I guess this depends a lot on what compiler/program is reporting the errors.
Good luck wading through whatever Google gives back for that.
Also I find they do well on error messages in general.
It's also far better than Google for recommending things (really anything) as you can use much more "precise" requirements for the recommendations and correct them, instead of getting at best some SEO spam.
It's obviously not good for stuff that requires real time updates like opening times, though I suspect with time OpenAI et al will combine their own search index + RAG to solve a lot of this.
The elephant in the room though is this is long term going to kill the incentive for a lot of content to be published. AI overviews have definitely reduced the amount of clicking through I do, even though I try and subconsciously 'support' sites by trying to find a link. But often it's right there.
I'm still going to turn to Google to find out what a store's opening hours or phone number is, as well as a lot of other tasks. But there are types of queries that are better suited for an LLM, that previously could only be done in a search engine.
There's also a non-technical reason for LLM search. Google built its business on free search, paid for by advertising, which seemed like a good idea in the early 2000's. A few decades later and we have a better appreciation for the value of the ad-driven business model. Right now, there's a whole lot of money being thrown at online LLMs, so for the most part they're not really doing ads yet. It's refreshing to make a query and not have sponsored results at the top of the list. Obviously, the free online LLM business model isn't going to last indefinitely. In the pretty near future, we'll either need to start paying a usage fee, or parse through advertisements delivered by LLMs as well. But it's nice while it lasts.
I think this is one of those points where LLMs have already changed the paradigm
People liked being able to search, but would not pay for it. For many queries , the value wasnt there: users still had to scroll thru pages or tinker for the right query for the required result.
Eventually search turned so much worse by seo spam, that kagi stepped in to fill the void.
LLMs start from a different direction. The value is clearly there. OpenAI etc still have a ton of paying subscribers.
I do think eventually some will incorporate ads, but I think innovation has revealed that theres a market -perhaps a substantial one- for fee-based information search with LLMS
But... your comment made me realize, a good use of LLMs with search would be not to ask them directly, but to use them as a front: detect what the user wants and devise the perfect Google search to find it.
LLMs should be very good at discriminating between a plain error message, a business name, or a general open question, building an excellent Google search for it, and parse and interpret the results.
It seems that's what Perplexity is doing, mostly. Personally I find it slower and more cumbersome than using Google Search directly, but maybe that's the direction we're all going. Machines to help us use machines.
Google abandoned the ability to use most search qualifiers effectively, so for niche websites I'll get zero results even with perfect exact-string-match queries, even when the site is in Google's index. On the other hand if I only vaguely remember the content and no related words or synonyms, Google is unable to turn my fuzzy feelings into an accurate search result. Plus, your ability to filter the "kind" of site is basically non-existent, making it all but impossible to find information on topics vaguely related to any word vaguely related a product you could potentially buy.
LLMs hallucinate and have their own set of problems, but that orthogonality makes them very useful situationally.
Not too long ago I needed to track down the blog Google references internally in the design doc of their TGIF employee voting platform, regarding Wilson scoring (using confidence intervals instead of means/medians/...). That's very easy to do in Google search if you can remember the right keywords from a decade ago (like Wilson scoring), but otherwise it's impossible. Reframing that problem for an LLM, you'd use plain English to describe everything you know, add a bit of flattery to shift the output distribution to that corner of the internet which actually knows what you're asking, potentially add one sentence to stop this last batch of models from wasting their time actually searching the web, and ask for a list of the top 5 authors and blog titles they think might be correct. That gives you a whole new set of search terms to finish your quest (in my case, the right answer was always in that list of 5, no matter how many times I reran the query).
That property of having to add additional context (e.g., when asking for recipes, I'll describe the background of who the LLM is roleplaying first) to get a good result is annoying. Full sentences with proper punctuation, unfortunately, also help. I wrote a small tool to make it easier for me to keep track of prompts I found useful and execute them with modifications.
As some other commenters mention, the LLM can help you with XY problems. You're searching for recipes, techniques, nutrition spreadsheets, ..., trying to craft something that meets some set of constraints (e.g., for some hypothetical set of guests you might require: no pork, most dishes have to be vegetarian, most dishes have to be gluten-free, it's fine if cooks all day so long as that isn't active prep time, you'd prefer to make it as tasty as possible while leaning in to cheaper, homestyle cooking, it has to use these red bell peppers I have already, and nutrition doesn't really matter). That's a nontrivial collection of tasks with a proper search engine, almost all of which the LLM is piss-poor at individually, unless you already have a good idea of the kind of dish you want to make. However, if you in two phases ask the LLM to brainstorm a list of 20 meal ideas and then expand on your favorite (that would be a decent time to add any modifications) then that collection of tasks gets done all at once. You _have_ to be able to look at the recipe and decide if it's any good or not, so a beginner probably shouldn't do that, but for everyone else it's a huge time saver.
I mentioned that Google sucks at filtering the "kind" of website you'd like to visit. LLMs handle that great. Like always, you have to be able to handle hallucinations (in this case, by just going to the results and checking if they're any good), but consider a prompt like the following:
> The web nowadays has tons of ad-infested, profit-driven, barely legible SEO drivel -- even in the top 100 search results and even from huge sites -- but all the old websites like Sheldon Brown's wealth of bicycle knowledge are still out there. List the top three old-web resources I'd want to read to learn about grafting apple trees.
When I ran that, I got one commercial result, one journal with a wealth of paywalled information, and one forum with a huge collection of free information about the particular species of Apple I'm interested in.
For "apple grafting" in particular, Google does just fine (a bit better arguably since the 1st Google result was the 2nd LLM result, and it's the only one I really cared about), but the more commercialized the knowledge you're looking for is the more the LLM shines out in comparison.
Well somebody has to be responsible for promoting the cancer which converted search boxes from logical set membership definition strings to "Smart™ boxes," designed to piss you off and then direct you to a sponsored product page.
So instead of getting a list of pages that contain that error message the LLM can evaluate which ones are closest to my configuration (are they the same OS, are variable parts of the error message the same or similar to mine, does the page have a solution or at least a discussion, or just one person posting the error with no other info).
Additionally search engines generally suck at negative phrases. For example I'm looking for some JS library that doesn't depend on React. Current search engines really suck at differentiating between "Foo for React" and "Foo, no React required". LLMs are pretty good at this.
That being said LLMs are also imperfect, so I wouldn't want the results that they reject to be completely hidden, but they can be ranked much lower. I think the biggest improvement would come from searches with many candidate pages (like looking for some library that fits my needs) rather than rare ones (like that error that has only been mentioned twice on the internet). But lots of times I am just looking for something with a very specific set of requirements that could be evaluated by an LLM (with pretty good accuracy). That can greatly improve my experience compared to getting all textually matching searches and doing the filtering myself.
Qwant used to pretend being a champion of privacy, a "french made" tech, but with a search engine mostly based on Bing, lobbying with Microsoft against our interests, and with a boss sucking as much public funding as possible to finance luxury HQ locations and lavish cars...
Qwant collapsed and was sold, but now it is just a mediocre ghost product trying to capitalize on a stained branding and on being a French/european alternative.
https://en.wikipedia.org/wiki/Quaero
https://www.heise.de/news/400-Millionen-Euro-fuer-europaeisc...
There are 2 clear things that is so common in Europe:
Politics injecting a shit load of public money thinking that if you give the money you will be able to reproduce American company success and co.
In the end, the money is wasted for their own interest by big groups, intermediaries, and opportunists. When the thing fails, it is the fault of no one, it was just "too hard" and "maybe the money budget was not enough for such a subject"
This is totally different to what leads innovation like Google, where you have doers that create something first and when they are able to show or convince that they have breakthrough, money will flow in by itself.
And at the beginning cash is used for brain and development instead of giving big salaries to top management and political friends.
The second thing that is usual is the pattern with this kind of projects:
- corporate sucks all the money
- responsibility is shared between multiple actors to spread the blame in case of problem.
- project fail and corporate give up. "Not their fault"
- one year later the initial hype subject is back on the table (European search engine sovereignty for ex) and politics announce that they will spend that much more money to resolve it
- same corporate vampires starts again from zero...
- and it fails the same in a loop
Because that's the desired feature, not a bug. The system works just as intended. None of the people at the helm actually believe they can create competitors to US giants, the scope is to funnel public tax money into the right politically connected private industry pockets.
It's just wealth redistribution with a veneer of "sovereignty", similar to large infrastructure projects, except a lot more profitable since more people understand physical infrastructure so it can more easily be scrutinized for corruption, but almost nobody understands IT infrastructure, so it can easily gamed as a bottomless pit for your tax euros that constantly fails in a loop while you socialize the losses and privatize the winnings.
The fervent wish to only let deserving people get grants results in a huge amount of box-ticking and self-promotion that (in my experience) seems to select for self-promotion parasites rather than people who want to make useful products and get rich from customers rather than government funds.
The problem with public funding is not the what but the to whom?
It works fine if you put the right person in charge.
However, there are very few signals to prevent the wrong person being put in charge, as it removes most considerations / incentivizes private industry uses. Which themselves are already tenuous!
EU needs something akin to Stripe Atlas, but that is not what the politicians want because they want EU to be manufacturing industry only, you can always import tech from other places … <shaking head emoji here>
It literally can't work at all. When was the last time you went and bought an Airbus? Airbus doesn't make consumer products. Passengers are the consumers flying inside them but they're not the ones buying them, it's the airlines who only have a monopoly of 2 global players to choose from in a highly regulated industry with expensive moats to enter meaning Airbus and Boeing don't really need to compete cut-throat.
Governments excel at building large infrastructure and defense companies like Airbus, Boeing, what have you, not at building consumer products at scale like Google, Apple, etc sine their success is dictated by the consumer spending preferences, not by requirements a government makes up.
Communist regimes did not make the best consumer products, the free market did.
What an insane comparison.
I understand that competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. Amateur search engines like Marginalia even run on consumer hardware. CommonCrawl text-only is ~100TB, and can fit on my home server.
Why is no company building their own search engine from scratch?
Running an index is an extremely profitable business, from multiple points of view (you can literally earn money, but also run ads, you get information you can sell, you can buy mindshare). Everybody is looking for indexes beyond Google and Bing, but there are none. If it really is as easy as indexing common crawl, then I think we'd have more indexes.
The problem of being able to provide fresh results is best solved by having different tiers of indices, one for frequently updating content, and one for slowly updating content with a weekly or monthly cadence.
You can get a long way by driving he frequently updating index via RSS feeds and social media firehoses to provide singnals for when to fetch new URLs.
This is too slow for a lot of the purposes people tend to use search engines for. I agree that you don't need to crawl everything every minute. My previous employer also crawled a large portion of the internet every month, but most of it didn't update between crawls.
1: https://www.iana.org/assignments/well-known-uris/well-known-...
Back before I found Kagi I used to use it everytime Google failed me.
So, yes, given he is the only one I know who manages this it isn't trivial.
But it clearly isn't impossible or that expensive either to run an index of the most useful and interesting parts of internet.
Google search has seamless integration with maps, with commercial directories, with translation, with their browser, with youtube, etc.
Even though there's more than a few queries they leave something to desire, the breadth of queries they can answer is very difficult to approach.
But for me, if I want to look up local restaurants, I go straight to Maps/Yelp/FourSquare(RIP). If I want to look up releases of a band, I go straight to musicbrainz. Info about Metal Band, straight to the Encyclopeadia Metallum. History/Facts, straight to Wikipedia. Recipes, straight to yummly. And so on. I rarely start my search with a general search engine.
And now with GPT, I doubt I even perform a single search on a general search engine (google, bind, DDG) even once a day.
https://help.kagi.com/kagi/search-details/search-sources.htm...
They source results from lots of places including Google. One way that you can confirm this is to search for something that only appears in a recent Reddit post. Google has done a deal with Reddit that they're the only company allowed to index Reddit since the summer.
DuckDuckGo gets no answers if you specify only results from the last week: https://duckduckgo.com/?q=caleb+williams+site%253Areddit.com...
Kagi is fine if you do the same: https://kagi.com/search?q=caleb+williams+site%3Areddit.com&d...
edit: I don't think this is a bad thing for Kagi. I'm a very happy subscriber, and it's nice for me that I still get results from Reddit. They're very useful!
When I search Kagi for "Hacker News", results start with this fine text:
65 relevant results in 1.09s. 47% unique Kagi results.
So, other indexes are fillers for Kagi's own index. They can't target their bots to places, because they don't have the users' search history. They can only organically grow and process what they indexed.The first result is almost assuredly the right one, but either they're ruling out a lot of pages as not-what-you-meant, or their index is really small.
Kagi reduces mental load by default, and this is a good thing.
Having a low number of results has a net benefit of lower cognitive load for me, so I like how Kagi returns less results, not more.
But "more results" is counted towards a new search, and I didn't know that. Thanks for pointing out.
The only way there is a chance for me to afford Kagi might be to buy "search credit" without a subscription and without minimum consumption. And then it would only be good if they allowed more than 1000 domain rules and showed more results (when available)
It's great that they are developing their own index, but I'm skeptical that it makes up more than a tiny fraction of what they can get from Google/Bing. DDG has been making similar claims for years but are still heavily reliant on Bing.
This isn't to knock on upstart search engines. I think that Google Search has declined massively over the past 5-10 years and I rarely use it. More competition is sorely needed, but we should be be clear eyed about the landscape.
That way users get tailored search without losing scale.
Could you please elaborate?
Cliqz was the first time for me that a Google alternative actually worked really well - and it, or now brave search, is what parent was asking for :)
> We expect that advertising funded search engines will be inherently biased towards the advertisers and away from the needs of the consumers.
Brave Search Premium hasn't been around nearly as long as their free tier serving ads, and I'm not confident this conflict of interest is gone.
Having independent indexes is a win regardless though.
[0] https://snap.stanford.edu/class/cs224w-readings/Brin98Anatom...
Your facebooks/twiters/instagram/stack overflow/reddit ... And they all have limited expensive api's, and have bulk scrapping detection. Sure you can clobber together something that will work for a while, but you can't runn a buissness on that.
Aditionaly most paywalled sites (like news) explicitly whitlist google and bing, and if someone cretes new site, they do the same. As an upstart you would have to reach out to them to get them to whitelist you. and you would need to do it not only in USA but globaly.
Anothe problem is cloudflare and other cdns/web firewalls, so even trying to index mom and pops blog site could be problematic. An d most of the mom and pop blogs are nowdays on som ploging platform that is just another silo.
Now that i think about it, cloudflare might be in a good position to do it.
The AI hype and scraping for content to feed the models have increased dificulty for anyone new to start new index.
The decentralized nature of the internet was amazing for businesses, and monopolization could ruin the space and slow innovation down significantly.
The legal concept of fair usage has and is being challenged, and will best tested in court. Is the Golden Age of Fair Use Over? Maybe [0].
[0] https://blog.mojeek.com/2024/05/is-the-golden-age-of-fair-us...
If the site is less aggressively blocking but only has a per-IP rate limit, buy a subscription to one of those VPNs (it doesn't matter if they're "actually secure" or not - you can borrow their IP addresses either way). If the site is extremely aggressive, you can outsource to the slightly grey market for residential proxy services - for fifty cents to several dollars per gigabyte, so make sure that fits in your business plan.
There's an upper bound to a website's aggressiveness in blocking, before they lose all their users, which tops out below how aggressive you can be in buying a whole bunch of SIM cards, pointing a directional antenna at McDonald's, or staying a night at every hotel in the area to learn their wi-fi passwords.
I am familiar with most of that, and there is a BIG difference between trying to find a workaround for one site, that you scrape ocasionaly, than to to find workaround for all of the sites.
Big sites will definitely put entire ISP's behind annoying capachas that are designed to stop exactly this (if you ever wonder why you sometimes get capatchas that seem slow to load, have long animations, or other annoying slow things, that is why etc.)
And once you start making enough money to employ all the people you need for doing that consistently, they will find a jurisdiction or 3 where they can sue you.
Also good luck finding residential/mobile ISP's that will stand by, and not try to throttle you after a while.
You definitively can get away with doing all of that for a while, but you absolutely can't build sustainable businesses on that.
Those example 3 bullet points of today's improved 2024 computing power you list isn't even enough to process Google's scale 14 years ago in 2010 when the search index was 100+ petabytes: https://googleblog.blogspot.com/2010/06/our-new-search-index...
A modest homegrown tech stack of 2024 can maybe compete with a smaller Google circa ~1998 but that thought experiment is handicapping Google's current state-of-the-art. Instead, we have to compare OSS-today vs Google-today. There's still a big delta gap between the 2024 OSS tech stack and Google's internal 2024 tech stack.
E.g. for all the billions Microsoft spent on Bing, there are still some queries that I noticed Google was better at. Google found more pages of obscure people I was researching (obituaries, etc). But Bing had the edge when I was looking up various court cases with docket #s. The internet is now so big that even billion dollar search engines can't get to all of it. Each has blindspots. I have to use both search engines every single day.
Most of Google's 100PB is picture and video. Filtering the spam and deduping the content helped Google reduce the ~50B page index in 2012 to ~<10B today.
I'm sure all the LLM providers are already considering this, but there's so much important information that is locked away in videos and pictures that isn't even obvious from a transcript or description.
> But what if I don't want to search Reddit, stack overflow, and blogs from the early 2000s
That is a strawman. There are huge numbers of websites (including authoritative ones like governments and universities) and a lot of content.
> There is an entire working generation that never heard a modem sound and has never even made a consideration for making sure their content is accessible in plaintext.
If they want video they will do the same as everyone else and search Youtube. Different niche.
> I'm sure all the LLM providers are already considering this, but there's so much important information that is locked away in videos and pictures that isn't even obvious from a transcript or description.
That is true, but if you are getting bad search results (and the market for other search engines are people who are not happy with Google and Bing results) that does not help much are you are not seeing the information you want anyway.
Ya know... a search engine that was limited to *.gov, *.edu and country equivalents (*.ac.uk, etc) would actually be pretty useful. Ok, I know you can do something like it with site: modifiers in the search, but if you know from the beginning you're never going to search the commercial internet you can bake that assumption into the design of your search engine in interesting ways.
And the spam problem goes away.
Hmm.
There is no open web anymore. Google killed it. There are probably fewer than 100k useful websites in the world now. Which is good for startups, because the problem is entirely tractable.
No matter what type of market analysis I do, I almost invariably find there's something different that say, the Koreans or the Europeans are using. The Yelp of Japan is Tabelog, the Ubereats of the UK is deliveroo, the facebook of russia is vk.ru etc.
That's really the beach head to capture - figure out what a "web region" is for a number of query use-cases and break in there.
If you went into Blockbusters then there was actually a small subset of the available videos to rent. Films that had been around for decades were not on the shelves yet garbage released very recently would be there in abundance. If you had an interest in film and, say, wanted to watch everything by Alfred Hitchcock, there would not be a copy of 'The Birds' there for you.
Or another analogy would be a big toy shop. If you grew up in a small town then the toy shop would not stock every LEGO set. You would expect the big toy shop in the big city would have the whole range, but, if you went there, you would just find what the small toy shop had but piled high, the full range still not available.
Record shops were the worst for this. The promise of Virgin Megastore and their like was always a bit of a let down with the local, independently owned record shop somehow having more product than the massive record shop.
Google is a bit like this with information. Youtube is even worse. I cottoned on to this with some testing on other people's devices. Not having Apple products, I wanted to test on old iPads, Macbooks and phones. For this I needed a little bit of help from neighbours and relatives. I already knew I had a bug to workaround, and that there was a tutorial on Youtube I needed to do a quick fix so I could test everything else. So this meant I had to open Youtube on different devices owned by very different people, with their logged in account.
I was very surprised to see that we all had very similar recommendations to what I could expect. I thought the elderly lady downstairs and my sister would have very different recommendations to myself, but they did not. I am sure the adverts would have been different, but I was only there to find a particular tutorial and not be nosy.
I am sure that Google have all this stuff cached at the 'edge', wherever the local copper meets the fibre optic. It is a model a bit like Blockbusters, but where you can get anything on special request, much like how you can order a book from a library for them to get it out of storage for you.
The logical conclusion of this is to have Google text search becoming more like an encyclopedia and dictionary of old, where 90% of what you want can be looked up in a relatively small body of words. I am okay with this, but I still want the special requests. There was merit in old-school Alta Vista searches where you could do what amounts to database queries with logical 'and's 'or's and the like.
The web was written in a very unstructured way, with WYSIWYG being the starting point, with nobody using content sectioning elements to scope headings to words. This mess suits Google as they can gatekeep search, since you need them to navigate a 'sea of divs'.
Really a nation such as France with a language to keep need to make information a public good with content structured and information indexed as a government priority. This immediately screams 'big brother', but it does not have to be like that. Google are not there to serve the customer, they only care about profits. They are not the defenders of democracy and free speech.
If a country such as France or even a country such as Sweden gets their act together and indexes their stuff in their language as a public good, they can export that knowhow to other language groups. It is ludicrous that we are leaving this up to the free market.
If you leave it up to the government, inevitably you're going to get only information approved by the people in power in that government.
You could call that search engine "Pravda".
(I don't think the open web is dead, but it's looking awfully unwell).
Excerpt from the article: >[...] Caffeine takes up nearly 100 million gigabytes of storage in one database and adds new information at a rate of hundreds of thousands of gigabytes per day. [...]
Your comment did make me pause and sanity check the math: https://www.google.com/search?q=%22100+million+gigabytes%22
In any case, a lot of people translated "100 million gigabytes" to "100 petabytes" based on that blog : https://www.google.com/search?q=google+search+index+estimate...
What's the current best estimate of its size now in 2024?
Another major factor is that building a search index and algorithms that searches across billions of pages with good enough latency is very hard. Easy enough for 10s of millions scale search but a different challenge for billions.
Some claim(ed) click-query data is needed at scale, and are hoping for that remedy. Our take is what is the point of replicating Google. Anyway, will this data be free or low cost? You know the answer.
Cloud infrastructure is very expensive. We save massively on costs by building our own servers, but that means capital outlay.
Are any individual users downloading CC for potential use in the future?
It may seem like a non-trivial task to process ~100TB at home today but in the future processing this amount of data will likely seem trivial. CC data is available for download to anyone but, to me, it appears only so-called "tech" companies and "researchers" are grabbing a copy.
Many years ago I began storing copies of the publicly-available com. and net. zonefiles from Verisign. At the time it was infeasible for me to try to serve multi-GB zonefiles on the local network at home. Today, it's feasible. And in the future it will be even easier.
NB. I am not employed by a so-called "tech" company. I store this data for personal, non-commercial use.
Finding some hack to democratize&decentralize the indexing and expensive processes like JavaScript interpretation, image interpretation, OCR, etc is an open angle and even an avenue for "Web3" to offload the cost. But you will ultimately want the core search and index on a tighter knit cluster (many computers physically close to one another for speed of light reasons, although you can have N of these clusters) for performance so it's a hard nut to crack for making something equitable for both the developers and any prospectors and safe from various takeovers and enshitifies. Let us know if you know a way.
Thanks to all the SPA idiocy you will miss enough content to matter if you have zero JS interpretation, so you would want to let the user choose which indexes they want for a query because sometimes you need these other resource types to answer the query.
Moreover, get out of your echo chamber and you’ll see that for a majority of humanity Google is the internet. You have to supplant the utility, not just the brand. Most businesses cannot handle the latter let alone the former. If you want something to replace Google, you have to think about replacing the internet itself. But not many are that bold.
why do you think amazon's trying to lay people off? they want them to create new startups that they can acquire.
1. Create your own index/crawl and true search engine on top of it, rather than delegating that back to bing or Google. In doing so, solve or address all the related problems like SEO/spam, and cover the “long tail” (expensive part)
2. Monetize it somehow to afford all of the personnel/hardware/software costs. Probably with your own version of Google’s ad products.
3. Solve user acquisition. Note that Google’s user acquisition process involves multiB/yr contracts with Apple and huge deals with other vendors who control the “top of funnel” like Mozilla Firefox, and that this is the sole purpose of chrome/android/chromebook/etc who you’ll never be able to make a deal with. You will probably at the very minimum need to implement your own platform (device, OS, browser alone probably won’t cut it).
4. Solve the bootstrap problem of getting enough users for people to care about letting you index their site, without initially being able to index a lot of sites, because you are not important enough.
5. Somehow pull all of this off in plain site of Google (this would take many years to build both technically and in terms of users) without them being able to properly fend you off
6. Somehow pull all of this off in spite of the web/web search seeming like it’s going to die off or fade into irrelevance
OR you can dedicate a decade of your life to something more likely to succeed.
There is one way to compete with Google search. Google search is a general search engine. But what if you only wanted medical information? information about cars? information about physics? electronics? history? etc.
Specializing makes for a much smaller scope of the problem, and being specialized means it could deliver more useful results.
For example, imdb.com. I don't ask google about movie stuff, I go to imdb.com because it specializes in movie info.
Some of them charge a lot for access but they certainly exist.
You run this bloated mass as a co-engine to your search. If you stumble upon any of the digested sites or articles you just have the site-eater blob regurgitate a html re-creation locally. They get no traffic, you get whatever they wrote purged of all links and with optional formatting to remove the fluff BS copywrite that many of these sites pad the tiny core of usefulness with.
A solution could be more than one algorithm being used to rank results, i.e. other engines, other rules. They'll likely use many of the signals available to them that Google uses for quality and relevance, but highly unlikely a genuine alternative search would rank them exactly the same- and much more unlikely an SEO could rank well in multiple engines.
The aff links aren't the problem, it's the proliferation of pages that are created solely to rank and get the links clicked on. Sometimes the content is useful, sometimes it's padded nonsense.
What do you mean? Google gained dominance just being a dot com URL. This in a time when competitors were already very well established.
Google's dominance will never be upended by an incremental improvement.
Only people I saw manually typing the url are my office colleagues(mostly engineers of some sort), even management does the same(skipping the FB part obviously) and you can see it when they are sharing their screen and they want to search for something.
To pull people out of google moat, first step is becoming their default search engine.
Eg see how chatgpt grew
Even then after using Google Gemini subscription the last few months I think the problem is not Google search rather the web ecosystem as if google gives me most of the answers without having to click any link you and I might be happy but billion people living off those links directly or indirectly won't be happy.
Here's one list:
>A look at search engines with their own indexes
https://seirdy.one/posts/2021/03/10/search-engines-with-own-...
HN discussions:
https://news.ycombinator.com/item?id=26429942
The only viable option when starting small is to grow in the cracks of Google with features that users really love that Google won’t provide.
Also important are user provided blacklists and whitelists of domains. Experts-Exchange was a site that called for this on search engines like Google, as Experts-Exchange was universally a cash sticky-trap waiting for searchers in pain to pony up. No, and constantly pasting -site:xyz each time is unacceptable.
ADDENDUM COPY&PASTE:
One blind spot I missed: advertising and telemetry activity are to count strongly in the commercial score.
When I slide the Commercial slider to zero, only web pages which there are no obvious detectable ways for the host or author to make money from showing me the page should be returned.
Bad behavior stems from the drive to make money. Make a search where it is impossible to make money by being a result in the search (especially if a Commercial slider is slid to low or zero).
Would be much simpler to calculate and detect. I would be super surprised if Google's algorithm doesn't already know how to detect affiliate links.
When I am looking up baking recipes or paleontology text on T-Rex, that should not be a commercial[izable] activity. Full stop! Commercial activity is like sexual activity in this specific regard: it's unwanted! 98% of my searches are for information that autists would pour their hearts into for free out of love, and that commercial interests will only detract from (such as pirated content and software cracks... lol wut did u xpect hm?). If I could pass a law which bled webmasters and search engines dry in fines until something like this is set up... I would! (Elect me BDFL of the world.)
It's downright silly to believe these creators won't behave rationally, given those incentives (eg. lean more and more toward pushing products with the highest affiliate payouts).
I've observed this myself in pretty much any niche I follow. Every youtuber/creator eventually devolves into an affiliate shill, no matter how honest and high quality their content is at the beginning.
If you're getting paid when you say something is good, you're not trustworthy. Period. What's funny is we all used to understand this! The internet broke everyone's brain.
Maximizing ad profit on search traffic necessitates more paid ads, both on search results page and the web pages. They won’t say that directly of course but that’s what we’ve seen play out over last 20 years.
Vague metrics like “good user experience” get trumped by hard metrics like total $ / search.
When I slide the Commercial slider to zero, only web pages which there are no obvious detectable ways for the host or author to make money from showing me the page should be returned.
Bad behavior stems from the drive to make money. Make a search where it is impossible to make money by being a result in the search (especially if a Commercial slider is slid to low or zero).
Instead of bending the web into something it is not, have you tried simply visiting a library and getting all information from there?
Thanks for describing the detection algorithm, I will now design my spam sites to defeat it.
This is how SEO works. There are no silver bullets.
“The only moral income is my income”
It’s definitely an interesting POV. A lot of the way Google works comes from the fact that people use Google at the moment they’ve already decided they want to buy something. Versus say Instagram where you might see ads but you haven’t decided to buy something and then use Instagram. It comes down to what people other than you are using search engines for.
You missed a key part of the comment:
> Offer the user a commercial activity slider. […] if a Commercial slider is slid to low or zero
It is exactly the point at which returning sponsored links maximally screws the user over.
- Has its own index
- Had a pro-privacy privacy policy well before DDG existed
- No conflated marketing (DDG wrt it passing on 3 octets of an IP to Bing for local searches, Ecosia for not including Bing's carbon footprint as a meta search engine)
- Has an API
It's a smaller index and has more limited resources, but pretty much the best genuine alternative, and great for finding older resources that have long since been buried by G/B.
Understandable knowing Ecosias goals, but I find it rather concerning their vision of a better search involves deciding what is good and bad.
Ranking by quality (against spam & SEO sites) is fine, but it should be applied equally to all Websites, and not target specific companies.
I too find this a bit strange. Downranking results that would otherwise be naturally highly ranked seems only to inhibit the operation of the search engine.
the entire purpose of a search engine is to do this, you've been grossly confused about the entire space if you think this isn't exactly what everyone is trying to do.
Although come to think of it I'm surprised that I haven't heard of any attempts of fundamentalists to make "moral" search engines that do things like exclude evolution.
Even after rephrasing things, more details, special quotations etc that everyone knows as the 'search tricks' the results are terrible.
Moreover, I can block sites and customize my own search results. This feels good.
When I first started using Kagi, it felt like leaving a closed building and stepping out to open air.
One of the few software tools I care to pay for except Jetbrains.
Only in practice I almost never use filter or block features, because Kagi does out-of-the-box what we always wanted to do our selves in Google: block spam sites.
The ranking also seems to be better for some reason somehow.
The funny thing is it doesn't feel like a step forward, but rather like a step back to Google ca 2009 - 2012 somewhere.
I honestly think google's monopoly on search at this time is 100% powered by momentum, there is almost no other reason to use it over something like DDG or hell, even Bing!
I have no knowledge of this field but something like that would seem seem to make sense.
Then things like whether the crawler should render the page (Using the end DOM content rather than the original source), does it do any tokenisation of the content, store other metrics etc, or does that need to be done by the end search engines.
Also there's issues with crawling Reddit, sites behind Cloudflare etc that others have went into more detail on this comment page.
Trying to reinvent Google/ Search in 2024 seems a bit like jumping the shark
Sure, garbage tier searches will be done LLM style. But few smart people might be bored with that, and pay for something better.
GNU/Linux as a desktop OS is an example of this market. 95%+ of people will work with Windows or MacOS their whole career which fits their use case perfectly. But the 5% of powerusers who choose Linux gain so much productivity and professional value that it's a thriving ecosystem with plenty of lucrative businesses built around it.
I want to browse a catalog of interesting web content.
Like walking down the isle of an old fashioned newsagent, flicking through the magazines that stand out.
For general information like recipes, tech stuff, how-to's, etc LLMs blow search engine results out of the water. Even the local/contextual content is better. The new maps feature in ChatGPT is amazing.
The only case where a search engine beats an LLM for me is when I'm looking for a site where the site itself provides a bespoke experience like shopping, where the engine provides a quicker path to the url I don't yet have in my browser history, like "<Some band> Tour" -> somebandofficial.com
Like many, I tried many other search engines, starting with DuckDuckGo way back when. I always ended up Googling (or !g… -ing).
Perplexity is the first one that consistently works for both code questions (what’s this error message) and local questions (where’s my nearest store X and when do they close). Now they just need to speed it up a bit - Google queries are effectively instant.
So, for me, "search" has fragmented into two separate use cases, one of which is served by the new not-really-a-search-engine.
I gave up on it because "two search tools instead of one" is just a really hard sell for me. When I go to search something, the keystrokes/hand motions are totally intuitive. Having two different patterns was annoying. StartPage works great for me. I have no reason to replace it, much less half-replace it with something that has only proven to be incrementally better at a few tasks. That's a matter of personal preference, though.
Kagi itself runs on top of Google. It's certainly worth paying for (IMO), but let's not kid ourselves -- it is *not* a competitor (as in, can replace Google). Without Google, there is no Kagi.
Google isn't the only search engine they use (they have a small selfmade index, Brave, Marginalia, and others which they don't disclose, and "vertical information" like Wikipedia, Wolfram Alpha, etc..), but probably most of the results are from Google. Without Google there would still be Kagi, but maybe not as good as today.
I was hoping Kagi got some media coverage and this was an article covering how it's so much better than Google.
My problem is with companies who start by declaring they are going to take on Google. These kind of companies are nowhere to be found a few years later.
It was a little odd at first but now I find I won't go back. I love I can ask follow up questions, and it shows the listing results on the right just like a search engine. But it also gives my a such summary on the left, again, with the ability to ask follow up questions.
Anyone trying to build a search engine to provide a Google-like experience is going to fail, in my opinion.
chatGPT search is a game changer. I'll pay $20 month for this forever.
We all hate how shitty the big sites have gotten in the name of maximizing profit, but I think it's more insidious than that. It would be one thing if running (say) a news site with decency and integrity was merely less profitable; there would still be people doing it. But I fear that it's actually become impossible to sustain a business like that. The ones that try either die or sell out to survive. (Or are so limited in scope that one person's unpaid part-time labor can sustain them.)
If someone is going to disrupt Google, it's because they've cut out the middleman that is search results and simply give you what you're asking for. ChatGPT and Perplexity are doing the best here so far afaik
Search engines are just cheaper to run. I don't know that there's a good, long term model for a free LLM-based search replacement because of how much higher the operating costs are, ad supported or not.
It could also become a differentiator allowing multiple suppliers. On one hand, you have people doing search for quality results. Other search engines include the AI results. The user could choose between them on a job by job basis or the search provider might, like !G in DDG, allow 3rd-party AI search as an option.
The bigger problem I have is with scale for the dollar. Search companies with their own indexes already mostly failed. There’s a few using Bing. It’s down to just three or four with their own index. Massive consolidation of power. If GPU’s and AI search cost massively more, wouldn’t that problem further increase?
To be clear, I'm not arguing for everything should be part of a pretrained LLM, but the experience of knowledge searching that ChatGPT and Perplexity provide are pretty superior to Google today (when they work).
If Google mangement can't see the user value of putting their own useful products in search, what hope is there for the rest of the world's useful information?
Another idea is to provide rating for some local products in exchange for searching.
I'm still wishing that one day we might get real microtransactions.
https://techcrunch.com/2024/11/11/ecosia-and-qwant-two-europ...
The adtech parts of ai are going to be weird.
NFL scores (live!) and random stats like point differentials or F1 driver / event queries. Of course the dozens of coding questions a day in this era where we all jump languages so frequently we have to keep looking up how to do basic stuff. Cooking instructions, and things like "can I substitute butter for lard". Product documentation stuff like "how can I assign a wheel on my Alesis QX61 to the pan in Logic Pro, and record it live to a track" or "how can I disable a malfunctioning pitch wheel on that controller". Historical and science queries. Language queries.
It has completely changed my usage, very much for the better. Web search provided RAG has been a game changer.
And while Google is outrageously fast, getting results on the screen quickly isn't the end. Then I'm digging through shitty bloated webpages or looking through a hundred Reddit comments filled with misinformation.
I just viewed the front page, looks like the definition of "internet chum".
But seriously, to upend Google, you are going to need to be the default on what people use, which I think for now is phones.
Another barrier, maybe they need to get away from this: "will require succeeding at home and growing revenue, which largely comes from running ads." So what do you do when everyone runs some ublock-origin thing? We need to figure out a search monetization beyond "feed me weird things I do not want" on the sidebar. Should websites pay? No, wait, then only those with real money are on the web. Should we pay? Now, wait, we have had it for "free" for too long (could be wrong, I might at this stage in the game pay to have an actual search engine, like say google circa-internet 2004, of course the net was a different place, but still).
While I think Google sucks right now and we need something new, this specific reason is so dumb. Unless I add the word "train" or "air" etc. I would much rather be shown either all options or the one that I care about most (if it's flying, then so be it - the search engine can't and shouldn't try to filter out options FOR me without my consent)
> Generative AI and new rules targeting tech giants are giving Ecosia and Qwant fuel to challenge Google and Microsoft and develop a web index for Europe.
“New rules” means regulation. “Market forces” kept the incumbents at the top, it is regulation that is giving others a chance.
> But Kroll believes tech advances have made affordable indexing more possible, and new EU regulations limiting the power of gatekeepers such as Google are making it a worthwhile pursuit.
Which links to another Wired story:
https://www.wired.com/story/europe-dma-breaking-open-big-tec...
They’re talking about the Digital Markets Act.
https://en.wikipedia.org/wiki/Digital_Markets_Act
Which is also what you’d get inundated by when doing the search I mentioned above. None of this is hard to find. If you want to discuss the article, it’s expected you make a minimum of effort to check what it says.
For a big site I help run, we’re getting about 8.2x the impressions on Google compared to Bing.
For you as a site owner, Google is the best: it delivers the impressions.
For a user who wants to search, Google has gone downhill since around 2009 and the only thing that confuse me is why DDG - who initially felt better - chose to run after Google down the path of insisting to give me results for things I didn't ask for.
(The usual answer is: "It is so much harder than in 2009 and SEO is so crazy these days, that nobody can do it, not even Google", to which I have to point out that even marginalia - run by one Swede - manages to do it in the niches it prioritizes.)
I have the same frustration. The killer feature that got me to switch from Google to DDG was that DDG would reliably return results for the search query I entered, long after Google had stopped doing so. Now that they've taken the same path the benefit is much less. Although I suspect this is more due to a change on the part of Bing than a conscious decision from DDG.
Censorship…for reasons. How’s this any better.