A quarter of the websites on the internet are behind Cloudlfare/Akamai with anti-scraping enabled. Of course they want to end up on Google, so Google/Bing get a green light from Cloudflare.
IMHO, that basically makes it impossible for anybody else to spin up an index and compete with Google Search. They cannot just scrape, they also have to work around anti-scraping measure by these CDNs. They have to maintain these workarounds, while trying to scrape as much as google has scraped so far, which is already a feat in itself.
Once you have this, you have everything: a critical mass of users, ads, data set for AI, ...
I dare anybody to scrape Reddit or Stackoverflow from scratch today.
According to estimates[1], Google indexes ~60 billions pages and Bing is only ~4 billions. If these numbers are close to reality, Bing index is 7% of Google's.
IMHO, this leads to all the problems people complain about Bing. The algorithm is shit because they have less data to fine tune it, people use it less because the algorithm is shit or the results don't show up, so they cannot collect data to fine tune the algorithm more, ...
When Google was just starting to gain popularity its search results weren't much better than Altavista's. Yet, its interface was much cleaner, it loaded and spewed out results faster, the overall experience was better. Bing is the Altavista. The user is bombarded with cookies, logotypes, ads, maps - all that complex layout right from the start. As a user you feel napalmed by information, and nobody likes that.
I do remember yahoo's home page was terrible, although I think there was a clean search.yahoo.com page no one knew about. Back then (and even now) having the search page load fast was pretty important so this always seemed like a nasty own goal.
Their analytics is present on many top sites so they can use browsing patterns to improve their results, on top of telemetry they get from Chrome. Perhaps only Facebook has something close to rivalling that kind of information.
Googlebot and to an extent Bingbot are granted access to most sites that want to be crawled. Any new entrants have to jump through lots of hoops (e.g. Cloudflare) simply to crawl the same web they do.
The UK's Competition and Market's Authority estimated that a credible competitor to Google would need to spend $20 billion just to be in a position to compete.
There's a couple of companies that could realistically compete with their financial clout, one of them is paid ~$20billion a year to be the search default on iOS.
For example, kaiser now references google.com instead of googletagmanager.com so people are tracked individually when handling their medical issues.
and they do it offline too.
In the early days it was search that drove chrome adoption and now its the other way around. With the low quality of search results, I doubt many would even notice if the default search engine is switched to something else.
They have over a quarter of a trillion USD in capital reserves and have acquired (and subsequently given up on) over 250 companies
They are in my view, Microsoft in its anti-trust era. This is only amplified by the more pervasive and aggressively anti-consumer industrial-user relations policy settings than the 1990s had
I should be able to move off everything else after I retire.