Is Google Getting Worse? A Longitudinal Investigation of SEO Spam in Search [pdf]
downloads.webis.de
downloads.webis.de
Surely people can relate to the situation where you end up on an article based on some technical query you have. The article repeats your question 7 times, has endless casually-related filler text that still does not answer the question and then ends with: try to unplug it.
It is so freaking obvious that it's a malicious content farm, but Google with all of its technical might seem unable or unwilling to detect it. If tech can't do it, organize some type of curation or feedback?
Same for image search. You search for "red flower Thailand" and flowers of various other colors from various locations appear. The idea that Google is spectacularly good at subject detection from imagery does not seem to actually work out in practice.
Most people's search queries consist of just 2-3 words. Nowadays Google consistently just drops the last word as if it knows better than I do what I need.
High value elaborate articles on various topics do not rank. Instead, dated articles do. You have to manually bookmark high quality content as you see it, because you'll never find it back via search.
Is everybody asleep at Google? This is not a small thing, this is your bread and butter. Teens are using Tiktok for search, you're in real trouble and better start cleaning up your act.
Now, if any of these flowers are next to a red dress, tapping the dress will reveal links to places you can buy it.
Google is not asleep. It has just got its priorities wrong. (Or rather, incentives in this organization seem to reward not what users like me appreciate.)
https://www.wired.com/story/prabhakar-raghavan-isnt-ceo-of-g...
And it's not just Google. Amazon is a mess. Social media is a mess. It all used to kind-of work but it's rapidly falling apart.
probably easy to confirm with a web crawler and generic searchs
it's really not a deeper technical problem.
Now I just need some kind of open source search engine to run on it ... (a bunch of text files that render to HTML, and ideally following the links 1 or 2 levels deep)
~20 years ago Google desktop search was a fantastic piece of software ... very fast and accurate on your local files. I don't think something like that exists now, and maybe never existed for Linux
Search engines are extremely modular and Unix-y. You have a bunch of indexed corpora and you intermingle them at ranking time, with respect to a query. But unfortunately there is no real incentive to provide something that has measurably good results and is also open to your own data and modifications
The incentive is to make a walled garden out of it
State of desktop search is very bad but this is what I use. It is of acceptable quality.
Here you go: https://yacy.net
And if I may expand a little, not just for crappy Google, I also create alternative local knowledge bases at work.
I can't find anything at work. Everything is spread out across chat, Wikis, SharePoint, email. All having different owners, content may at any time disappear or move, there's constant authorization headaches.
Whenever I come across something useful that I expect to be of some future use, I make a local copy. File, web page, wiki, anything. Because our information systems are a massive failure.
There use to be a time when paid placement was only 1-2 results.
It’s frequent now that the top 5-6 results are paid placement.
(And when I’m doing a search for a specific product I know I want, competitors are bidding up those search terms which is annoying because I’m being shown not what I’m explicitly searching for)
it's largely just a product of ourouboros
So I would suggest that Google knows what it's doing, it just makes them money.
Google has commercialized a huge amount of search terms. I'll use biology as an example. You search for particular species and search results prioritize products that kill the species. You search for a particular plant and you'll have a hard time learning about the species as it only shows cultivated versions and products related to how to care for them.
Pure information/knowledge for the sake of learning and curiosity is de-prioritized.
Back to image search, it's unable to figure out original sources or doesn't care. Pinterest is the well known manifestation of that.
Google shopping results: completely broken. Click through on the products and half the time the price, availability, discounts and stock do not match.
Everything is so goddamn broken, and nobody at Google seems to care. I can't explain it, but it's been going on for a good 6-7 years or so.
I've seen this in many search terms. Purposefully push down the things that people are looking for and they will click on ads.
Those sites that have a ton of extra info? Many of them also sell in content ads right?
The ad machine took over long ago and you are forced to play the ad auction game if you want to show up in results, and if people are looking for something they are forced to scroll past a lot of ads.
I watched this happen in chunks over time. It wasn't always like this, but it is now, and it's been a slow creep over time.
Especially since you get the clear spam sites that somehow reference your query in the page content (where they've just spammed loads of keywords, but also pretty sure some spam sites are doing something dynamic with it).
We're all technically minded here but very few people really understand how technical choices add up to greater detriments.
and that's today's Google. they minimized the index and maximized the searches that yield profit through Google ads. those websites you hate? they monetize Google ad words.
Looks mostly red to me. A little pink too I guess.
Thanks!
I think this is an excellent methodology for testing the quality of search results. I would love to see a standard search engine test and scoring system based on this, maybe similar to some of the LLM scoring systems.
That approach also misses all the copied-a-github-issue low-effort content that seem to crop up on Google.
Affiliate links create misaligned incentives between content creators and consumers. This includes on the products listed themselves, where the affiliate is incentivized to select the products that give the most kickback rather than the ‘best’ for the consumer. But it also includes the content itself. Affiliate links create incentive to (a) write reviews, where a site wouldn’t have bothered before; (b) churn out lots of content with lots of links, that can be picked up by search engines; and (c) to not invest much in the actual reviews, and instead generate quick, low-quality content, since the content creator doesn’t actually care about finding the “best” product (which takes lots of time and money to do correctly), only on having readers click their links.
This is why sites with names like Celeb Rumor Central have dozens of articles like “top 10 coffee makers 2024”. They just hire a freelancer to do a few quick Google searches for coffee makers, then churn out 10,000 word articles with sections like “the history of coffee” and “why do people drink coffee?” Then the actual review is “this machine is purported to boil water and drip it through coffee grounds, and has high ratings on Amazon (we may earn a small commission when you click a link on our site)”. Increasingly, they’re just using AI to do it, cutting out the cost of a freelancer.
In my personal opinion, affiliate marketing is one of the worst things to happen to the modern web, and the source of a ton of content farm SEO spam pages.
Back when I was still involved in acquisition marketing, I did a test where I killed all paid search and built relationships with the affiliates to negotiate our "ranking" higher up the page. It was a huge boom for business. We scaled paid acquisition at a profitable CAC, which was very difficult to maintain, let alone scale, in bidding directly and it was significantly less work to manage.
Consumers value this type of content quite a bit, even if many are skeptical of the quality. Sometimes it's nice to just see a pared down list of things with even a cursory rundown of features/differences.
I've seen so many "Best $WHATEVER of $YEAR" and it's really "Best $WHATEVER $PRODUCED_BY_WEIRDLY_NAME_FLY_BY_NIGHT_COMPANY_ON_AMAZON $YEAR"
In practice, that is never the case. At best you're getting reviews that only compare products available in the same marketplace (say, for example, all the table saws you can buy on Amazon). At worst, you're getting reviews from vendors who offer the highest payout.
Also, a heck of a lot of consumer products are absolute garbage. (IMHO, most of them, but others may feel differently.) Who is going to write an honest, scathing review of a product and then monetize the link to it? Why even bother?
Almost? There doesn't need to be anything nefarious going on, but human beings align to incentives, basically always. So the trick with reviews is going to be to avoid any incentives coming from the product side, and embrace those coming from the consumer side.
1. Domain Interception & HTTP redirects 2. Tracking codes embedded in the URL directly
> Kagi surfaces shopping results featuring unbiased reviews and no affiliate links to help you identify the best product across categories. Top results include discussions focused on helping you find the best item to purchase - you are not bombarded with affiliate links and ads. Continue to scroll and you will see product comparisons across multiple vendors so you can pick what best suites you. Kagi's shopping search will always return a detailed discussion of which product to buy not a competition amongst advertisers promoting where you should buy. Kagi is focused on providing you the best results to make an informed decision not polluted by affiliate links and advertisements.
https://i.imgur.com/vPRpBYJ.mp4 (yes the recording is very broken but you can still see)
When I hover over the "i" icon it says:
> Information provided by Looria, Amazon, Reddit and other sources. Shop confidently with Kagi - our results have no affiliate links.
The sites featured in the "Shopping" widget were:
- RTINGS: has 8 affiliate links (just from one article I picked, it was featured several times)
- NYTimes: has 47 affiliate links (just from one article I picked, it was featured several times)
- Techgearlab: has 9 affiliate links
Even if you ignored the widget, the first result was a site with 7 amazon affiliate links.
A pattern I’ve noticed from people not working in tech, or more junior people in tech, is that they have a thinking that software can ever be perfect. If either of those categories describe you, I can tell you from experience no software is ever perfect!
One of these is "ads/trackers". I imagine that it would be feasible for this to include some of the more common affiliate URL types, or third party lead/affiliate tracking bounce hops like awin.
Clearly there will always be some amount of ability to "defeat" this kind of measure by obfuscating links, but eventually the user needs forwarded to a link that has a referral parameter or a site that sets an affiliate cookie or similar.
The "tracker category" also can give a bit of extra information - things like "invasive fingerprinting, advertising"
In addition:
* it demotes sites with popups (think newsletter sign ups)
* it demotes sites that block (or complain about) ad blockers
* it demotes sites with a high number of ads and favors sites with no ads
* it demotes sites using certain sketchy ad companies.
* It demotes sites that have paywalls
* It detects possible link networks and flags them for human review/removal.
* sites with RSS feeds get promoted.
* There is a toggle to hide all sites with ads or external trackers, but it is still WIP (The whole project is).
There are many other features. No idea if I am going to make it public, I created it to update my skillset. I actually thought about setting up a nonprofit and making it open source, but I haven’t decided.
I don't know if discovery is actually a bottleneck to be automated away. It might be the fun part. I'm thinking back to the Napster approach where you could browse other people's libraries for music ideas.
If I do an image search for the word 'strawberry', how many of those results are not stock images, images from a store, etc. of a strawberry? can you find an actual picture of a strawberry sitting in the wild? or just some picture of a strawberry a person uploaded without trying to sell you something?
Edit: This got downvoted to hell, so let me be more explicit. This study did not look at Google results, the title is pure clickbait. They used Startpage results as a proxy for Google results. I don't think that's a valid assumption, even if Startpage is using Google's index.
If the authors have done their due diligence and confirmed the results from Startpage are actually Google results, then I don't see why they couldn't claim so in their title.
It's pretty obvious why the results won't be the same: the feature sets of the search engines are different despite being based on the same index. Startpage's own documentation even has a page on how and why the results are different!
(Personalization is a part of the Google feature set, and has to be taken into account when considering the results. It's also a part of the Bing feature set. This didn't stop them from reporting the results of their testing on Bing in detail.)
The reason we used Startpage is simply that it's much easier to scrape. We started off checking only Startpage (as proxy for Google) and DuckDuckGo (as proxy for Bing), since they are both simple to scrape and produce stable rankings. Plain Bing, on the other hand, is a lot trickier. You often get different results for the same query and you get blocked much more easily if you send too many. The only stable way to use Bing is via the (certainly not very cheap) API, though even that wouldn't necessarily guarantee the same user experience as the web frontend.
We did notice, however, that DDG (despite being mostly Bing) did deviate quite a bit in their results, so we started scraping Bing as well for a fairer comparison. As for Startpage, we did check it initially and we found the results to be virtually identical, except for a few minor rank differences here and there (which are probably just geo personalisation). The differences may have become larger now that Startpage also taps Bing to a certain amount, though when I do spot checks, the results are still sufficiently similar and they are also sufficiently different from what actual Bing gives you. Most prominently, Startpage/Google give you a lot more YouTube results, which we did another small spin-off study on (to be published at CHIIR this year https://downloads.webis.de/publications/papers/bevendorff_20...). Moreover, we could also measure certain immediate effects of Google's ranker updates in Startpage, which weren't as apparent in Bing. So we are confident that Startpage is a reasonable proxy for Google, though it's certainly something we will keep in mind for follow-up studies.
I'm sure that using these sites was the thing for your research for all kinds of practical reasons, and you don't need to justify that :) My actual complaint was about the misleading title, which you did not address.
Note that the clickbait title actually hijacked your research; there's like three comments out of 240 here that are about your paper. So while it probably feels good to have received all this attention, it's just bogus engagement. I guess it's pretty meta for a paper about search result spam to do blackhat social media engagement optimization though.
But the worst part is, Google SEO has infected the entire web and made it into complete garbage. Hopefully, this last decade or so will just be a blip before we return to baseline, where it can be wild and free again.
* Of course, the stats should include the total amount of internet users globally, or normalize the amount of searches based on that...
When I first started out all the veterans of SEO kept telling me not to do this, don't do that with things that could get your site buried in the SERP's. At the time Google's algorithm was really good at ferreting out affiliate links, link farms and other nefarious black hat techniques SEO's used to game Google.
Now? Complete opposite. I have several freelance clients and I've used every dirty SEO trick in the book and all of them have worked like magic to get my clients sites ranked on page 1 or 2 of the SERP's.
I have no idea what changed, but Google is super easy to manipulate now to get your site or specific pages ranking really high. I haven't heard or seen any of the horror stories I read and people blogged about constantly when I first started out for years - which tells me they're all probably doing the same thing I am and not seeing any repercussions.
Maybe Google doesn't care because users have become so savvy, they can filter through a ton of garbage in minutes to find what they really want?
It turns out they appear to have a folder for basically every suburb in the country, thousands of subdirectories yielding the same looking page with that suburb name inserted into the text, making them look local no matter where you are. I can similarly find this with many other industries.
This sort of thing twenty years ago was a huge no no. Google talked about detecting mostly duplicate text and would bury you for it. Like a very basic rule was that if two pages were mostly the same only one would really rank.
Nowadays the "lead funnel" is an entire industry on websites following this pattern, and it clearly works.
It's especially interesting since you mentioned normalizing searches by the number of internet users. The country with the largest number of internet users is China, with more than 1 billion of them. And they don't have access to Google. And their local copycat, Baidu, is years behind Google in terms of technological sophistication and simultaneously years ahead of Google in terms of user hostility. So what do Internet users in China do in a post-search world? They simply open various apps and use the full text search feature of different apps. For general knowledge they might open ZhiHu and search there; for something resembling the old-time personal blogs by individual users they might open XiaoHongShu and search there; for short videos they might open Douyin and for long ones Bilibili. For reaching an organization be it a store or a museum or a hospital or a government department they might open WeChat and search there for an official account or mini program (a mini program is a website that uses WeChat APIs and can only be opened in WeChat).
I made these observations on a recent trip to China and it's clear to me what a post-search world looks like because China is already there.
Baidu search was fine during early days, issue as you hinted was PRC internet went mobile first and content got locked behind various platforms and made deliberately hard to scrape - crawling/indexing got locked much earlier than west. Hence now as more gets locked in west behind logins, western behaviours also shifting towards that model, how much of default search is query + wiki/reddit/youtube or straight into short video services like looking up recipes on XiaoHongShu. Reddit especially, simply because reddit app has horrible search experience. Also technically, Baidu rankdex predated Google PageRank which Larry Page referenced for Pagerank patents. Eitherway, depending on how ChatGPT copyright drama plays out, imo more people will just go the lazy route and have AI generate good enough summaries for most queries.
You are talking like open web is dead but it's not. There are millions of blogs and personal sites out there. Walled gardens are user hostile and hungry for money, that's why enshittification[0] happens.
Quick: If you want to search on what is the best toaster to buy, what URL do you type?
How do you find out the weather?
How do you find a local brewery in your area?
For me the first answer is Amazon. And the second is ask Siri. The third is Apple Maps.
Google is now below 50% of my search terms. Heck, I’m more likely to search Reddit for some queries.
Could you give some examples of search queries that would benefit from filtering by reddit?
(My own example: I've been looking for recommendations for a solid Linux laptop. A good result would be a list of reviews written from personal experience of owning such laptops. Reddit was useless for that.)
People actually discussing how they did this is the only real answer, generic websites with pictures and descriptions of his weapons never help.
I'm so happy that their new pricing scheme has done away with search quotas. Their claim that the average user does 100 searches a month seems absurdly low. I would blow through my 300 search quota in a few days.
Pay with Bitcoin [1].
If Kagi supported USDC I would definitely use it.
I’m not letting perfect be the enemy of good.
If what you mean is you don't believe they are not connected to your account, I can point you to the FAQ entry on their website.
https://help.kagi.com/kagi/faq/faq.html#what-data-does-kagi-...
I could take that list, as-is and use it as a block list on my entire network.
One interesting solution to the problem is to have more than one dominant search engine and its algorithmic choices, having half a dozen web-scale engines with some variation at least gives the user a choice into other avenues of information discovery. (There isn't really much point in using Startpage and DDG here since they're effectively meta search engines of Google and Bing). For SEOs in English speaking countries there is not much point in thinking beyond pleasing Google.
Clearly AI and whack-a-mole spam sites have been a problem for a while due to the prevalance of people tacking on 'reddit' to their query to find other humans talking about stuff.
I spent an afternoon Googling every possible incantation only to get useless AI generated text, travel agency sites or simply unrelated content.
I was about to accept my loss when I tried Kagi. The first page showed an exchange that accepted the currency. Very far from me and with terrible rates, but still.
Anecdata and all, but the fact is that I'm using Kagi more and more and it's winning my trust and good will fast.
Even as a happy customer, I can see how suspect this praise can be, but remember that HN is exactly who Kagi was created for. And I don't remember having spend $10 a month on something that had that much impact on my daily life. It's _just_ a search engine, but I do about 50 search a day, and having great results each time while the other one is giving me more and more garbage results is huge. So I completely understand why people satisfied with Kagi can be noisy about how great they think it is.
I think 2005-2010 Google was peak web search. It would have given me an obscure blog with the top 5 currency exchange offices. Kagi gave me one single crappy result, but it was still, sadly, better than what 2024 Google can do.
Google search for topics I'm unfamiliar with/wanting to learn about all lead to low quality, SEO-optimized to hell, clik-baity sites that are just riddled with ads. I have to add "reddit" to most searches just to find semi-relevant content.
But Google search for topics i'm super familiar with and just need a transactional search to look something up tend to be much better and generally the fastest way to accomplish a task.
BTW, maybe someone wants to create a very simple webpage with a search mask that allows adding a few (customizable) terms and options and simply forwards that to Google's search when pressing enter.
The second type of searches you describe seem to be better so far, but I've stumbled upon a bunch of obviously generated garbage recently. So not sure how long it'll hold.
I've tried things like DDG, YaCy, Bing and others, but often Google is just significantly better (but not necessarily good).
In longitude, latitude, and by many other measures :)
Plenty of earlier posts and comments about that on HN, for many years now. What's so surprising or new about that, then?
As the saying goes, it's news if a man bites a dog, but not the other way around - doggone it if I know why, man ...
This is particularly egregious with Python, and I suppose it must be just as bad or worse in the JS ecosystem.
I am able to run it in attick on raspberry pi. We do not have to rely so heavily on google.
https://github.com/rumca-js/Django-link-archive
It is true that it does not serve me as google, or kagi replacement. It is a very nice addition though.
With a little bit off determination I do not have to be so dependent on google.
Here is also a dump of known domains. Some are personal.
https://github.com/rumca-js/Internet-Places-Database
...and my bookmarks
https://github.com/rumca-js/RSS-Link-Database
Some more years, and google can go to hell.
but the structure where the enemies are constantly optimizing against your fundamental goal.
Like companies avoiding taxes.1. Spam is more than in past. The outrage porn, clickbait headlines etc. are lot more than in past.
2. Dominance of few domains despite poor quality content. For lot of coding related queries, dev.to, hashnode etc. appear in top results despite being clearly spammy.
3. Paywalled content. Most irritating part is sites like medium which appear in top results, have high value content and yet are behind paywall.
Internet is growing and so are Google's problems but I think they are still on top of things.
Except, if you put the same search in Startpage and Google, you get different results. Image results especially are quite different. Text results were mostly just a reorganization on my quick tests. (Tried the title of the paper as its own search "Is Google Getting Worse? A Longitudinal Investigation of SEO Spam in Search")
Edit: One other notable result is from Figure 3, that a huge percentage of the results now are Amazon and Youtube. Many orders of magnitude in most cases. Amazon (3000-4000), NYTimes (1000), Walmart (~500?), Insider/PCMag/Tomsguide (~50?)
> Startpage delivers Google search results via our proprietary personal data protection technology.
Are you allowed to search the internet with an ad blocker installed?
Do you use a special search interface that doesn't return results with ads?
eg just yesterday - a search for the NYSE listed company "betmgm" on google news [US] yielded 100 spam results [affiliate offers + bonus codes for BETMGM sportsbook]- and not one real non-commercial, non-spam news post concerning the company.
[data: https://gamblingindustrynews.com/news/affiliates/google-news... ]
Stay tuned here on HN.
What if Google flipped its SEO weights from positive to negative?
There's a thing that's been happening the last year or two where compromised web servers serve the intended content unless there's a referrer of Google (and maybe Bing?). The site looks normal to the owner, or to crawlers, but when a user clicks through from Google it redirects them to a spam site.
I mostly see it with restaurants getting hijacked by herbal quackery.
(By comparison, Phind gave the correct answer, and high quality sources.)
Looking to browse works by a specific artist? Good luck, Google evidently can't tell the difference between a genuine J. C. Leyendecker piece and anything shat out by a "Leyendecker style" image generation model. Search for "Yoji Shinkawa" and this https://i.imgur.com/RYghaoY.png is currently the first thing you see, which isn't even close to his style, but Google has somehow determined it's the image that best represents him. The full images page shows his actual work but interspersed about half-and-half with AI imitations.
My speculation is that Google prioritizes showing recent results, presumed to be the freshest most up-to-date information, but of course for a historical event like Tiananmen Square or an artist who died in 1951, nearly all of the fresh results past a certain point are AI simulacra.
Or they were at least threatened with legal action. I remember when it happened many of my friends were annoyed overnight
So rather than just delist Getty and solve the problem, they decided to make their product objectively worse
1. Domain-specific engines is too complicated. Let's have a general purpose search engine.
2. Too many buttons is too complicated. Let's have a single search box.
3. Guessing what you want when you are anonymous is too complicated. We need your search history to tailor your search for you.
Please, search engines, hear me: you can not simplify the needs and wants of 7 billion people and zettabytes of online content into a single little shitty search box! It's too complicated!
At this point the only thing these search engines are good for is for finding the URL of something you already know the name of. It's a phone book. I type Hacker News on google, it tells me its phone number: https://news.ycombinator.com/
Finding my preferred UI https://hckrnews.com/ that's not even in the first page of results. Rajat's version https://hckrnws.com/... nowhere to be seen.
Google is becoming a mediocre phone book.
What other buttons do you want?
But for sure Google has crippled its core product. First they removed the + option, which forced the inclusion of words (their excuse was that it "interfered" with their stupid "Google Plus" product which is now gone). Yes you can use "allintext:" but come on. I'm not even sure that's honored anyway.
And the removal of the ability to exclude certain sites from results.
Are you saying that's how Yahoo worked? IIRC, that was optional, they always had the search box:
https://web.archive.org/web/19961023235123/http://www10.yaho...
OK, I just looked up their old layout, and they buried the search bar amongst so much crap that it wasn't clear whether it pertained to the ad or the thing above it, or what... if you even noticed it: https://helios-i.mashable.com/imagery/longforms/04ILIeAX3JAF...
It will never work. It can't work. Anyone who thinks this can work is delusional. And with Google it's particularly obvious how stupid the trend is.
Look at the "tabs" Google has for search: all, images, videos, shopping, news, etc. This is things users can input. But wait! What if an user wants news about something? And they have to reach all the way out to the news tab. That's too much for our bubbling moronic users to manage! They can't into computers. They have room temperature IQ. They have never used Google before, so they don't know where the tab is. They probably don't know what tabs or links are either. I know what I'll do. I'll put an AI to reorder the TABS of my users based on their search history, query input, season of the year, and their zodiac sign based on what birth day they used when they signed up for a Google account. That should solve it.
And now the order of the tabs is all over the place and when you want to click the "images" tab it's sometimes not the second tab and when you want to click the "videos" tab it's sometimes not the third tab.
I think this is very interesting because you have to think. If Google can fail this hard at tabs, which is not really a complicated thing to program, imagine how hard they are failing at indexing the entire interweb. Imagine if they are doing to search results the same nonsense bullshit they are doing to the tabs. Just imagine it. It's clear they have absolutely no idea what they're doing with the tabs.
I also noticed this; they redesigned their search tabs few months ago and now it's sort of bad and user unfriendly. Sometimes there is "News" tab, sometimes there is no news tab and tabs look all the same and generic(white rectangles with black text).
>I think this is very interesting because you have to think. If Google can fail this hard at tabs, which is not really a complicated thing to program, imagine how hard they are failing at indexing the entire interweb. Imagine if they are doing to search results the same nonsense bullshit they are doing to the tabs.
God knows what is in their index and what is not in their index. I think every search engine needs to make its index open and transparent. And yea Google can fail with all sort of things, they are not ubermensch or something like that.
Google's ranking algorithms and search technology are millions of LOC and not even Google engineers know how exactly Google Search works.
Another search to try is "hand reference", tons of AI garbage
> like when searching for the "tank man" prioritized showing a fake AI selfie from the mans perspective
In this case the label is that it's a AI-generated and the source is an article talking about that AI-generated picture.
They own the search market anyway, and the more time you waste on their platform searching for what you want to find, the more ads you see and the more money Google makes.
What else are those 97% people gonna do, "Google on Bing" instead?
So for them, being bad is actually more profitable than being good, meaning there's a conflict of interest between what Google provides and what their users want, but since there's basically no equivalent competition, they get away with it laughing all the way to the bank.
That's not the only way - it's in their best interest to rank ad-laden pages higher. That's more ad-impressions, which is how they really make their money.
Lets say a user searches for $FOO. Why on earth would google return the most relevant result if that result is ad-free? They can return the second-most relevant result, and get impressions on both the search-result page and the page that the user sees when they navigate to the first result.
Searches that end you up on AdSense (or pages that could have AdSense) probably are much less likely to have intent. So instead they pass you off to something you will read so they can hit you with retargeted ads from the last time you searched with the intent of buying you something.
I am sure there's overlap, but probably not a lot.
When you have +95% of the market, you can merely be good enough that users don't leave.
In this case, until Google literally starts serving ads-only and links to other google products (Hi Chrome!), they are not going to lose any money.
Bing's AI shows some promise, but I can't find anything better than Google. DuckDuckGo is shit. Classic Bing is shit. Brave Search showed promise, but it's also full of spam and with a smaller index. Marginalia showed some promise for smaller websites, but it's small, too.
All of them are unusable for local searches, except for Google, which is where I need search most. That and searching for obscure programming-related error messages.
Kagi doesn't even use it's own index, you're basically paying for a UI making API calls to Bing and Google. How well does changing the ranking work anyway, if you don't have your own index?
I want to believe that if Google were more aligned with users, search would be better. But where is that better search engine to showcase it?
Do a poll on HN. I bet that the vast majority of users here still use Google's search, and you can't blame HN users of not knowing of alternatives.
If Google really keeps its market share due to their monopoly, where is that better option that's being ignored? Tell me and I'll jump on it.
Although Moore's law brought down the price of hardware and information processing dramatically in the last 20 years, it is still fairly expensive to crawl the Web, index it and rank it. Hardware cost + engineering cost can escalate pretty quickly, unless you decide to have smaller index than major search engines but then users will complain that search results are not good enough.
Due to the defaults it's the only search engine most people ever heard of. Think of Plato's cave allegory.
If all your life you've only used Google and never anything else, making it the ground truth for you, how would you know it's bad in order to motivate you to look elsewhere?
LLMs
I decided to first try ask an LLM, so I asked Bard some targeted questions and asked for sources, and had all of the answers I needed, conveniently bullet pointed within 3-4 minutes, and all I had to do was go and verify the sources, job done.
My favorite is recipes.
The blog spam around recipes is notoriously bad and is a direct result of Google Search.
Using these LLMs, all I need to do is tell it what I'm looking for in general and it will give me an entire recipe with no extraneous information.
It can handle adding or removing ingredients, substitutions. It can adjust servings. It can flip the recipe to work in a slow cooker or pressure cooker.
I've even had some limited success where I list restrictions based on picky eaters in the house for creating longer meal plans. No search engine can compete with that.
There were apps designed around these use cases. I foresee these sort of nuanced and personalized interactions being a key to drawing people away from search engines.
I'm going to give it a go for recipes, because as you say, blog spam is a nightmare. My wife has a lot of cookery books which she likes for her style of cooking, but I'm a little more "ad-hoc".
I haven't used ChatGPT extensively so I can't make a comparison, but with the right questions, I've been able to get great answers to technical things I can't be bothered to look up by hand.
Now I am using 3 search engines.
Bing for Normal searches. Google for local(country) searches. Yandex for small sites, blogs, forums.
Edit: They aren't likely to add it. https://kagifeedback.org/d/687-implement-regional-pricing
Google knows so much about me. And yet it continues to act as though it doesn't.
I initially trusted Google for its efficient and seemingly fair service, where smart ad-targeting was the price for speed. But now, Google feels similar to Facebook; it's harder to switch to alternatives like Kagi on iPhone due to financial ties with Apple.
This shift in Google's approach, prioritizing trapping attention over genuine service, is frustrating. I'd rather pay for a search service that values my time and provides real utility, than endure the hidden costs of 'free' services.
Increasingly you have to pick a destination and blindly follow directions, your ability to use it as an informational tool, but exercise your own judgement seems to have been intentionally crippled.
How much would you pay for Google search without ads? Or more importantly, how much would the average Google user be willing to pay?
If you live in the US and click 3-4 ads per month, you're generating ~$10-20/mo in revenue for Google from advertisers.
(Like you, I would also love to pay for an ad free Google, but sadly advertisers are willing to pay Google a lot more money than consumers are willing to)
Today? Nah. Search has become bad enough that I don't believe it's worth any cash.
Due to the constant trash these algorithms insist on feeding me, which I believe is actively contributing to the dumbing-down of society, I've naturally gravitated towards books. Books are leagues better in terms of quality of content. They are more detailed and thorough. It's a richer experience so far.
I see a strong metaphor for literature authored before a specific time, roughly when the web came to be, or certainly when two way discourse on a page, or aggregation, became prevalent. And certainly far before bespoke communications targeted at us, as individuals or interest groups.
If the modern web were air it would taste of metal; I have a fear that I will become biased that older texts are superior for the sole reason that modern texts can be assumed inferior.
This part really worries me right here. We don't know the long term ramifications of the modern day internet, but so far they do not look good.
Even if Google were incentivized to clean up their algorithms, the people who want to make money will necessarily always be one step ahead in the cat-and-mouse.
First, they have been injecting unrelated recommendations into search results, under headers like "For you" and "People also watched". I don't want to see a pimple-popping video when I'm searching for something related to woodworking. (That's an extreme example, but I have actually seen those types of videos injected into completely unrelated search results. In fact, I just did a search for "hand cut dovetails" as a test and YT recommended two different disgusting pimple/pore videos in the first few dozen results.)
Second, they don't admit that they are out of results. Instead, when search results dry up they coughing up unrelated recommendations so that you can keep scrolling forever. This makes it look like there are more results than there actually are, which is completely unhelpful.
YouTube actually has a lot of decent content still, if you can find it. I've found that jumping over to look at channels that do collabs with channels that you already know is a good way to discover new content.
I don't know about other topics, but it should still be pretty easy to find quality content if you want to learn guitar. Spend a couple of hours to find the channel that suits you best and stick to it.
Overall, there's still lots of great content on youtube, but you need to look it up yourself. The recommendations are useless and make you lose time, and the suppression of the dislike count makes it much harder to filter bad stuff. Also most professional YouTubers don't have much to say. They repeat themselves over and over.
Maybe someone should come up with a custom recommendation algorithm. Don't know if that's doable.
I watch only certain channels and a lot of long form documentary and gaming content. So that is what I get recommended. The few times I use the "dont recommend" button it never pops up again.
It also is helpful to search for videos and click on those. If I search for long form podcasts or interviews, I will get more of those kinds of videos recommended to me.
https://addons.mozilla.org/en-US/firefox/addon/hide-youtube-...
works.
Recently I was thinking about getting a dog and was trying to find as much information about certain breeds as I could. Everything was just the same stuff repeated over and over, very little substance.
It took me a few days to realise but a lot of the Youtube results were actually entirely created by AI. Not just the voice over and what it was saying but the actual video was as well.
You can get AI to spew out content on whatever topic and as long as there's money to be made from said content, this does not bode well at all.
Repeating a comment I read here a while back - Google has lost, perhaps even given up, the fight against spam content.
The more market share of online display advertising they gained, the worse their results got.
Why? Because the only way to have a large volume of authentic content being produced at scale is to have a healthy ecosystem of independent sites that are profitable based on display ads.
As much as HN-types hate advertising, it was literally the only thing that made the web of yesteryear so special. Things like Adsense enabled blogs on tons of niche topics to be monetized and thus we had better open web content.
When Google decided there was more money to be made off ads before you even clicked on a search result, that ironically was what ended up killing search.
Now the only way to monetize content from search is via shill company blogs, affiliate marketing listicles (10 best dog toys), etc. So that’s what we get.
For example, if there was a passionate person creating authentic, amazing content about dogs, they wouldn’t even crack page 1 on any search for dog toys no matter how good their content is.
So basically the problem of search engines and Google in particular is discovery. All early Google adapters say that they loved Google because it gave them relevant results and because they discovered new websites on Google. Nowadays they "discover" SEO spam, ads and the usual suspects like the most popular sites in that search category.
That's why we need "discovery engine" for Web, something like TikTok but adapted to Web. If search engines were the evolution of web directories, we need to think about how we can evolve search engines.
Like, literally irrelevant. I'm still watching YouTube, using Gmail, and occasionally checking out something on Google Shopping if I want to find something locally instead of on Amazon. But I use Google search about 90% less now than I did a year ago.