Google search's death by a thousand cuts
matt-rickard.com
matt-rickard.com
Where in the past I could find documentation, forum-posts, wikis - concise information - these days I find SEO-optimised marketing fluff, link-farms, clickbait, long-form articles with thousands of words that ultimately say nothing, political tirades that don't make sense for any non-U.S.-citizen, and ads - tons and tons of ads.
It would seem that it was Google ads and Google search's own algorithm - which directly led to this state of affairs.
And that's the main reason why asking some LLM for answers is so much more attractive than using search: the AI cuts through all that crap, and extracts the actual information. If it's not hallucinating, that is.
The web is no longer for humans or for finding information. It's for bots, crawlers, for the Google algorithm and for marketing.
Humans simply can no longer digest it - we need AI to generate something more concise, more to the point, more digestible. Without the marketing, fluff and bait and filler.
I mean, if they can make billions by serving 99% of people, why would they spend billions on making the remaining 1% happy?
But we've been on a race to the bottom in most sectors for a long time now.
Google makes billions from advertisers - the companies who pay for ads, the companies who create those SEO-optimized marketing-pages. Those are the target audience, that Google wants to make happy.
Why would they care about us users? We don't pay for search.
It's not just the quality of websites that is negatively hit either. The real victims are us, civil society and our future.
The information realm that is the internet should provide not only some basic tidbits, it should inform to the full extent of human knowledge. Even more than that, it should provide us with the freedom to discuss among ourselves and reflect on ourselves, humans and humanity, about the future of society.
Without such self-reflection, there is only stagnation and retardation. We cannot change what we do not understand and we understand ourselves the least.
Plus we've all seen what this has lead to in meat-space. Information is one sided. BBC is supposed to be the paragon of neutrality, for example, and while it is pretty neutral on the situation in many places, it's not exactly neutral on UK and Australia political issues for example. And it suffers from the same problem as pretty much every other meat-space information source outside of libraries: the only information you can get from it is really quite superfluous.
Their recent "is AI evil?" series, for example, is neither neutral nor does it give enough information for people, even if someone wanted to go really in-depth on the issue, and have a really well-considered opinion, the only place where you could collect enough information to do that would be Youtube imho. With some blogs and maybe a scientific paper or two. Oh wait 3 out of those three work with a lot of help from Google (just try finding scientific papers without Google scholar)
I agree the situation is bad. But frankly, while Google perhaps has some blame in this, I'm very sure the situation would be 10x worse without Google.
You put up a false dichotomy, as if Google disappearing would mean a void to wander the world eternal.
The internet is very much like public infrastructure. It is crucial to explicitly determine what functions there are indispensable for society to thrive.
Ironically, finding available information is just such an indispensable utility. Imagine roads only being present or usable when conducive to advertisement.
There's also a big interest in tracking people for reason du jour.
I usually hear this "freedom of speech" stance until somebody hears something they don't like, then it's out the door and censorship gets a foot in the door.
As long as there is no consensus in society what speech is ok and what's not, we won't get anywhere. Currently, we are mainly leaving this to platforms to decide, and they cater to their customers (which are ad clients), not to what is best for society. Obviously there needs to be some moderation and banned contents. But a line needs to be drawn and anything beyond that needs to be free. No matter if some individual gets hurt by it or not.
I suspect that thanks to android phones collecting all of our offline data, and chrome collecting everyone's browsing history, Google no longer needs our web searches to collect the most intimate details of our lives which is why they aren't interested in spending time and money on their search product. In fact, the more useless google search is the more google can make companies pay for ads placed at the top of listings because if those companies don't pay up, their customers may struggle to find their site at all through all the spam.
The (not golden) goose lays golden eggs. The recurring time-based nature of the eggs is the key aspect of the fable.
It would only provide blogspamish recipes for canning your own cabbage soup linked through the tiled suggestion interface. Not a single text-style link was returned and none of the results were pre-canned soups. I scanned though page after page without getting anything else.
Oddly, when I searched "canned cabbage soup history" it switched modes and started providing normal results, including the product and Wikipedia.
What strikes me about this is that I was actively searching for a product I wanted to buy. If there's any business case for ruining Google with weird specialty searches it would be selling products.
Why is there a "provide blogspam" mode?
Also, I'm now asked if I'm a robot at least five times a day.
I actually thought some of the “google is dying” handwringing was overwrought until I experienced this. It was just, bad.
Because blogspam sites all implement Google Adwords for monetisation, meaning Google can now show you 30 ads instead of just 3 at the top of the search result.
Remember, Googles customers are the advertisers and the site-owners implementing Google Adwords, not you.
You're the product being sold.
If I do scroll through the search results, select a few promising ones, wait for the page to load (slowed down by multiple fat ads), click away the cookie warnings, reject the GDPR consent, disagree with the newsletter popup,switch to reader mode to remove the visual clutter, scroll down the page skimming the content for the right paragraphs... then I will eventually get the same information, that ChatGPT could just have provided me instantly without all the hustle.
Rather than gate keeping, LLMs actually compress the search index into a manageable size, enabling easier sharing. While it's impossible to download the entirety of Google, you can conveniently download a LLaMA, promoting both privacy and autonomy. This marks a departure from the days of having every click and search keyword tracked.
Perhaps the advent of LLMs and Generative Image models by 2030 could allow us to cut the cord and navigate a virtual internet, have a room of our own to dream and be free, essentially functioning as an augmented imagination where we set all the rules.
Local LLMs solve the old problem of accessing information without leaking out what you are looking at, something that bothered me every since BitTorrent and copyright wars.
The side effect is that the below-median brainpower of the 30% of the audience who doesn't know any better than to watch ads and click on them is now the driving metric for content.
In a world where programming revenue comes from ads, and only idiots see ads, the programming will cater evermore to the tastes of the most-idiotic viewer.
This is what I don't want from a search engine. What I want from a search engine is a list of links to places that have relevant information. I don't want the engine to summarize or distill what those places are saying. Too much knowledge and understanding is lost by not doing that myself.
I think one of the fears some have about AI/machine learning/etc is that they realize their entire advertising & marketing world could become irrelevant. Users could get short, accurate answers without having advertisements in the way.
Kind of like how things were many years ago.
Also hate the top 10 pages whenever you search the best something, like best domain name registrar. I don't want to read a spam blog post with affiliate links, I wish google would show me the actual domain registrars instead (like chatgpt does when i ask it). Google has been gamed so badly and they have been doing nothing about it just because the spam blog posts contains their ads.
There are some tricks I learned on HN to use uBlock origin to filter these spam sites but Google really needs to fix this. There is only so much an adblocker can do to fix search. And right now all the useful content is getting blocked while the spam content is not only allowed but ranking on top of everything.
When gaming search engines became a profession, the end of search appeared on the horizon. Guess we're headed back to web rings and link indexes (which will be consolidated, heavily monetized, and abandoned). If we're lucky, we'll be back to dialup BBSes by the 2040s.
It really seems high-quality search is fundamentally in opposition to serving ads, alas. (At least once every page in existence probably serves ads via the search operator's network.)
I have a lot of complaints about, say, McDonald's, but they make their money through giving people something of value to them. Advertising legitimizes the making of money in a way unrelated to value delivery. (The same is true about a lot of finance.)
When you combine that with up-and-to-the-right numerical goals and standard executive incentives, over time you pretty much guarantee what Doctorow calls "enshittification". Delivering value becomes at best a side effect of the system.
"[W]e expect that advertising funded search engines will be inherently biased towards the advertisers and away from the needs of the consumers"
and
"[W]e believe the issue of advertising causes enough mixed incentives that it is crucial to have a competitive search engine that is transparent and in the academic realm."
Both from The Anatomy of a Large-Scale Hypertextual Web Search Engine by Sergey Brin and Lawrence Page (1998)
... but more seriously: Yeah, their reasoning (now) isn't exactly going to be unbiased and/or unblemished by their own experience. They probably have the most extreme survivorship bias ever. (Not their fault, they just do because of whatever factors got them into the position they're in.)
There are, of course, a few relatively successful paid general purpose search engines but these serve a niche demographic if you consider the world-wide scale of google et al. Possibly specialized search (we build one) will be able to thrive in the future, but these engines also serve niche markets in the end.
Thus, the real competition to ad-based search is not high quality search and that is likely why search results don't get better.
Google should have leaned into just this small segment of tooling and done much more of a ban hammer on the bad actors. They didn't and a bunch of people have left. I use DuckDuckGo mostly and sometimes Google but never by default. I don't even like DDG that much but it's good enough.
The popular "Awesome ${whatever}" collections are very much that, link indexes for relatively narrow areas.
Curation matters. This is something computers are still bad at, or are too expensive to deploy for a mass-market service.
The challenge will be staying true to not showing ads, respecting user privacy, and not requiring a subscription. So far, the only thing that works is free daily quota + pre-paid
Now we can use ChatGPT to filter through Google's infested mess, but this double edged Sword of Damocles will be able to create infinite attempts to bury genuine content with ad spam.
Like in the back of a book, or an old subject based card catalogue ... Dewy-decimalise the web :)
I've always felt that an index by subject would be more useful than string-match based searching. Of course, the index might rank links within each sub-sub--sub-sub...category with something like the original page-rank.
Now if Yahoo (or whoever) could avoid the enshitification trap ... imagine what a fabulous resource that could be.
Now that is an awesome idea. I really do think we need to have category-based indexing as well as page search.
Do you mean self-managed?
Everything else is effectively the influencer scene. Which is increasingly deplorable as well.
Anything with wide enough reach becomes cost-effective for gaming.
So one would have to return to a highly fragmented world to make gaming the system cost prohibitive.
And that would get us to a pre-Internet world. But then again, it’s not entirely unthinkable that we’re headed towards increasing Internet fragmentation if various governments get their way.
That amount of money is probably more than Google now makes from my online presence because I adblock, block 3rd party cookies, tend to click "block" to everything including the idiotic "legitimate concern" and never ever click on ads.
wow, Search really isn't worth developing.
Edit: not beam but color
Fuck it, I'd pay $15 a year to have a Google search that puts as much effort into finding me the shit I actually want as Google does today in serving me BS ads I never pay attention to.
They announced "Google Contributor" at one point but it never went anywhere
The ads industry would not like their reach to be limited to those not paying for premium search.
Tell me exactly how much the advertiser paid for his placement, and that's a hugely important signal here.
If I'm searching for weird hobby parts, even though it's a high purchase intent query, they're probably paying pennies per click.
But if you start searching, say, financial stuff and the ad placement figures start showing multiple dollars per click, it's a warning "these people are willing to spend THIS MUCH MONEY to present a message to you, this probably means there's something sketchy involved."
I know, for example, anything pertaining to insurance and financial products is highly likely to turn into a farm of cross-selling and personal-information harvesting, because the cost per acquisition is so high and the tendency for everyone to sell the information to everyone else is so great.
Would 3 EUR or even 10 EUR / year for each customer be enough to run operations?
Of course people game the system, apply shady practices, sell courses with tips. It is a nuance topic, the name now is bastardized by marketing companies doing shady things, but on his core is just practices to create good websites and provide a good user experience
Unfortunately, there's no going back - the closest we can get is BBS-via-SSH. The entire landline phone infrastructure is crumbling around us (or in many cases, completely gone). Voice calls are packet-switched now, rather than circuit-switched as in the past. The upshot is, fancy modulation techniques that made full-duplex 33.6k possible over voice-grade connections aren't going to work, and even good old Bell 103 (300 baud) may end up being problematic.
I'm not sure I can even get new landline phone service, and if so, it's going to be expensive - and the wire plant is an unmaintained mess. When I got my folks off their landline and onto VoIP some years back, their old landline had so much hum it was nearly unusable. Once the inside wiring was disconnected from the landline and connected to an ATA, the hum was gone. It wasn't our wiring.
How many of those thousand domains will result in billion dollar fines?
The point is, downranking decisions by google have a non-trivial chance of being litigated in court, and at that point, I doubt that the decision will actually be made on the merits of the particular site. A sympathetic website owner with a sympathetic regulator is frequently going to win against google, regardless of how shady the site actually is.
first ten results include: theworkathomewife dot com › google-ads-quality-raterHow to Get Paid to Be a Google Ads Quality Rater
Sep 18, 2022 · As a quality rater, you ensure ads on Google – and other search engines – are both accurate and visually appealing. You will be given sample search terms and potential
Forbes - After Months Of Protest, Google Search Quality Raters Finally Get A Raise
ARS - 2017 - https://arstechnica.com/features/2017/04/the-secret-lives-of...
and the google note:https://support.google.com/websearch/answer/9281931?hl=en
with a hand wavy - these raters don't actually effect the results..
I read dozens of pages of the raters guidelines pdf years ago - and I can see how they are being sneaky with things the 'down rank things that lack trust signals' - which in of itself sounds okay - especially with some situations.. but that exact thing can be used to push down entertainment sites, and can help adwords make more money by forcing lots of sites to pay to show up in the top 10 or miss 90% of the internet traffic.
So I agree that they are pushing shitty results in many search verticals - but yeah it does seem they are using thousands of people to create them this way on purpose.
It's no longer a better search engine - it once was - today it's popular because it's a default on many devices, and the network effects make it so people build their sites and update their listings like it's the only yellow pages in town.
Their high-brow censoring leaves much entertainment and other things better found using other portals to search and find imho.
Oh I did use ublacklist which mixes blocking in both of the google search and the google image search. But curious what is the filter you have in your uBlock origin?
One of the results was "title (recommended)".
That should be enough for Google to ban the result.
I'm pretty sure SEO sites will be able to figure out don't title your page "title (recommended)".
Then websites started to optimize for what Google actually parsed. Meta tags, no things behind JavaScript, semantic markup etc. And that worked really well. Stuff was easy to find.
Those were the best days. There was some luring, but at least if you were looking for something technical you found it (provided it existed and was indexed).
This is no longer the case. Give you an example:
Any search term for Windows+Error+KB<number> is a nuclear wasteland of companies having copy-paste pages only to sell you their services at the end. Lately the same has been popping up on YouTube.
Any search for an issue with formatting a harddrive lists 20 pages which sell formatting tools.
A search for checking whether a certain animal is dangerous yields wildly inaccurate results just to sell me pest control. Spoiler: it was a Jerusalem Bug. You don't get a fever at all. The thing just has huge mouthparts, and that can physically hurt.
When you put "(recommended)" in your title because you know it's then displayed verbatim in Google is nasty.
All driven by these websites which 'recommend' things to you.
Kagi recently released a leaderboard for per-domain customization: https://i.imgur.com/ViLamx7.png
They are 2nd for "lowered" and first for "blocked".
Add Twitter to that list and I'll start finding accessible content again.
If I'm not on a functionally unlimited unlimited plan (rate limiting is fine if it's reasonable) then I have to think about it every time I search. Should I really be searching here or should I be using the search box on stackoverflow or github or Wikipedia? Just introduces some constan cognitive frisson to the experience.
Personally, my record was below 900 searches and usually below 800, so the 1000 searches (technically 1500 because early adopter) are absolutely fine for me.
More investment in Blogger with better linking between blogs (the way you can @someone on Twitter) and discovery (#tag searches allowing you to see posts from multiple blogs, curated blogrolls), more realtime/live logging functionality, easy microblogging functionality, etc. would probably have led to a lot more content existing on Google owned spaces, as opposed to being siloed behind Facebook and Twitter paywalls.
Because they kept chasing "the next big thing" aka "the next product to give me promotion inside the company".
Remember Google Wave? Knoll? Plus? Orkut? Spaces? iGoogle and Sites? Slide? Jaiku? Buzz? Aardvark?...
(I had to scroll through https://killedbygoogle.com to remember some of them)
It blocks copycats and hide them from multiple search engines. You may also use the list with uBlacklist.
Absolutely. I have a browser extension installed specifically to block domains that are this type of spam from my Google search results, and I’m adding to it almost weekly.
Come on Google, you’re seriously ranking that as front page worthy?
I recommend this userscript https://github.com/vladgba/Back2source for avoiding Stackexchange clones, it saved me a lot of time.
Google stated goal was the "organize the world's information". Nevertheless, they didn't come up with Wikipedia, the highest quality curated human-readable information repository. They also didn't come up with ChatGPT, arguably the first LLM good enough that can perform non trivial tasks of data recall and organization.
Google had early success with search. Then decided to cement their lead by just throwing money at "all the smartest people"[1] so that their competitors would have difficulty hiring. This put Google in the situation where it's better to spend its resources fighting to keep the advantage that they have (by controlling the Internet infrastructure; android, chrome, email, etc) than innovating.
Google is dead man walking, unless something very substantial changes.
[1] "all the smartest people", read, "academically successful but rule-obeying, unimaginative, and risk-averse people".
Things are somewhat different today then when Schumpeter was writing. On the one hand, governments and venture capital firms are now major sources of R&D funding. On the other hand, the growth of the financial sector has to some degree crowded out higher-risk investments within large firms.
My take is that it was taken by its own weight. The government fines & incorporation that started to fear those fines slowed down the progress. How difficult is to launch a change in that environment? Same for the baseline -- people react to changes. You consider the baseline to be the Golden standard. How difficult is to launch a radical change considering that?
IMO, the only way out is to shrink the product. Split, shard, whatever. Be small enough compared to the rest so you are not the default target (Coca Cola, Procter & Gamble, etc. do it right).
Microsoft and OpenAI nailed this in my opinion. The funding deal between the two was genius. OpenAI gets billions of dollars from Microsoft that it then spends on Azure.
Microsoft got priority access to ChatGPT, equity in OpenAI, high-scale testing of Azure, plus some margin on OpenAI's usage for training. At the same time, it _didn't_ get the legal liability of running ChatGPT. Maybe it could eventually be hit for allowing training models with copyright data on its servers, but that feels unlikely.
It got all of the access to cutting edge productized LLMs that it could have wanted, including financial upside, without the bad parts.
Any high churn change is difficult to get out at that scale. Now add all corp-fear because you are under lense of all governments. You can't keep the startup mentality with that kind of environment.
Wikipedia was good before the culture wars got to it.
Not that anyone asked, but I think he is absolutely guilty of spreading vaccine propaganda - however, I think one editor has a very good point about how this should be approached. Here is what they said:
>I too would avoid the pejorative word "propaganda" on WP:TONE grounds. If they must be included, terms like "propaganda" and "conspiracy theories" should not be stated in wikivoice. Instead, it should be stated dispassionately what the sources say. See Deepak Chopra for some examples of more appropriate wording: "His discussions of quantum healing have been characterised as technobabble"; "The ideas Chopra promotes have regularly been criticized by medical and scientific professionals as pseudoscience." (emphasis mine). HappyWanderer15 (talk) 12:46, 17 June 2023 (UTC)
I'm still curious how they're going to properly monetise ChatGPT at scale. I can't wait for the AI ads...
Right now they are producing value for consumers, in order to attract investors. Investors will want a return: the product will be the consumers.
Not only did Google make the main breakthroughs that all LLMs still rely on today (mainly attention architecture, and large scale training of DNNs), but they had a chatbot similar to ChatGPT (Meena [1]) way before ChatGPT was a thing.
Granted: they didn't release it to the public, so there is a failure of innovation there, just not on the technical/machine learning side.
[1] https://ai.googleblog.com/2020/01/towards-conversational-age...
YouTube is just bollocks now too, it contains a stunning amount of junk content and clickbait.
If someone tells me they learned something from YouTube, I actively question their authority on the matter at all. I didn't feel this way 2 years ago. YT seemed to be a generally credible source of information. Words cannot describe how polluted my suggestions are. I actually find it disturbing how much fake, incorrect and bias information is being suggest to me all the time. If I took it all on face value I'd be an idiot.
I've noticed recently that even people I thought were pretty credible have upped the clickbait severely in recent times. I guess this is what it takes to get seen but it's at the point where I just don't open YouTube unless I already know what I'm looking for so I can avoid it.
I have some doubts about the quality of Wikipedia content. Everything that can be politicised is politicised , everything that can be played in order for someone to advance his/her plans or point of view is played.
Apart from bad faith, some article lack relevant info or contain errors due the volunteer nature of it.
And if it's not - then what's better?
I don't see a way to access my local university library for that data in the same way Wikipedia is available.
Wikipedia suffers from massive bias and subjectivity issues, and also from accuracy problems. Yet, all the attempts to do better fizzled out so far, and most non-niche knowledge and information repositories have the same or worse issues.
If we look back at history, many of the earliest written works were published in this way, as a philosophical debate between different fictional characters.
However it seems increasingly difficult to find good online debate with several sides of an issue. People tend to silo themselves into their own bubbles, and denounce any platform who even dares to let an opponent speak his mind. I remember just a decade ago where you'd find all opinions on the spectrum on the same message boards. Now, people prefer to shelter themselves instead.
Beyond search, they are controlling the Android ecosystem. Your predictions may be immature.
Oh yeah we are really seeing all those successful forks taking off. /s
On a more serious note, the key parts are not open source, that's why all those alternative ROMs are having a hard time and nobody can really run Android without official builds outside a niche of enthusiasts who are fine with sacrificing certain functionality.
The main reason competitors don't take off (e.g. Tizen) is the chicken/egg problem of the Google Play Store (which can't be shipped without Android).
Interestingly, when I tried this just now, I did get the paper as a result. I could swear I tried this same search yesterday and had zero results.
Not sure if some are removed also but it definitely changes.
as someone who remembers How Google Used to Work, Yandex feels like Old Google. It indexes the hell out of long tails, will find exact, specific terms and phrases, and 100% respects generic search operators such as exclude "-".
do I look for news articles on it? hell no. do I find the unfindable (on google, and increasingly, on duckduckgo)? yes, I do.
When I'm looking for a technical solution I combine words. But Google gives me the middle finger, returns results for the first one which is like 'Windows Service' only to show me a generic website which has a download link for every KB number out there.
This makes non technical question-answering easier, to the chagrin of every techie.
I think your problem is a mix between a lot of things, but mostly it seems to be SEO farming.
So I'd imagine the OP's query should work when the image id is wrapped in double quotes. Otherwise, it'd mean that it's not in the index anymore.
It's like trying to write a Gmail filter with the "=" and then a word wrapped in quotes. You would THINK that would be a strict X = Y... but it's not.
At some point a few years ago the Gmail team made the "=" operator a "fuzzy equals", so for example it matches on parts of a word, even if you explicitly say you want to filter a word with a space at the end because it's "close enough".
Drives me fucking mental - I have an email filter that instantly archives every email from Jira except those that have my user name in the body of the notification (meaning someone has tagged me).
In Outlook I simple filter for "my_username " (with a space at the end), and it's able to tell the difference between that and my email address. In Gmail, "my_username " somehow matches both "my_username" and "my_username@email.com"... meaning EVERY FUCKING EMAIL I GET MATCHES THE CRITERIA.
Even worse than giving no results is giving you page after page of "results" containing slightly different numbers/codes.
With BERT they can have a model to embed all the search queries for 99.99% languages and works sort of fine for 99% of the queries, so they are fine with that as well.
It would have been sort of easy detecting that a query works better with the previous index than the new one, but it seems that they decided against that. So we lost a lot of usefulness, like phone number search.
Same result: Now it's nothing, or unrelated SEO.
A week ago, a search for < James Palmer manchester foreign policy > wouldn't return https://foreignpolicy.com/2017/05/23/i-love-manchester-but-p... on the first page while < James Palmer manchester site:foreignpolicy.com > would [1] [2]. This appears to be fixed as of today but this kind of kindergarten level search intent not being fulfilled would have been unthinkable for Google even just a few years ago.
[1] https://twitter.com/BeijingPalmer/status/1670904508191322112
[2] https://twitter.com/shalmanese/status/1670973880637493249
Still waiting on an explanation of fix for this one.
My assumption would be that the search results are garbage because the query is extremely ambiguous in the search engine's eyes, especially if you omit or modify any of the words (which it is allowed to do because it is not a search for an exact phrase).
So many people and groups want to profit with as little effort as possible that the commons (freely available data, open APIs, FLOSS software, etc.) is being overrun with little to no regard for the long term effects.
Data can be copied endlessly, but it has to be generated the first time and updated or else it becomes stale and its usefulness decays. If no one is willing or able to generate or update it, there is nothing good and the signal-to-noise ratio falls off sharply. Everything is noise and it sucks.
A lot of technology and culture was built without such “protection”.
I sometimes wonder how recent history would have played out without copyrights and patents.
Case in point: look how far LLM's have come since llama hit the public sphere vs. how far organisations like google managed to get, with more time and more resources at their disposal.
Now that we're all equals, the rich and powerful don't have to do any such things anymore.
Don't know if they let search rot or they broke it intentionally. But it got broken regardless of Reddit. Reddit has a lot of info I guess, but it is far from being the only website on the internet. However seemingly google stopped indexing more than the top 150-200 websites on the internet (and even those results are often lacking the searched words).
I'm just trying to wrap my head around it but I think I'm making sense of it now and it's kinda like that race to the bottom I see sometimes mentioned here? Do I got it right?
On top of this, siloing adds the extra bonus that it breaks things like The Internet Archive. There are a lot of communities on Slack and Discord that are just prime for vaporising without a trace. A TON of knowledge will just disappear once someone decides to turn it off, for whatever reason.
My suspicion is that this was the beginning of the public presentation of corporate infighting, where what the public saw from Google began to be less about good products and more about the products that particular execs, managers, and developers needed to "win" in order to bolster their careers (or what they perceive as such). This is why Google has developed a reputation for killing useful products, often replacing them with flashier and more frustrating alternatives (if at all).
I don't like putting this on external factors. I think Google is doing this to themselves. It's a more private and stretched-out version of what's happening very publicly with Twitter right now.
And surprisingly it's free.
The end result is companies trying to capture more of that value for themselves and not letting Google eat their lunch, which required moving content and communities behind closed doors.
I suspect that the independent internet died for pretty much the same reason why many people won't post anything interesting or authentic on social media anymore: Being active online can lead indirectly to real world harm to yourself or your family. Evil people use the internet to harass other people and to attack their livelihood or put the safety of their family at risk. Many employers now discriminate against employees or even just people applying for jobs who openly express opinions on the internet. And even posting pseudonymously isn't completely safe because it is usually possible to deanonymize posters if an attacker is obsessed enough.
The sad reality is that many people do inauthentic things for the sake of career advancement (think of all the college students who do extracurriculars or projects for the sole purpose of "resume padding") and refuse to do things they actually believe in for the sake of their career. When creating quality online content becomes a risk to career advancement, people are going to stop doing it. In the long term, that means a siloed internet full of corporate propaganda, SEO optimized blogspam and clearly inauthentic posts.
In short, I believe Google search is getting worse because the internet is getting worse. And the internet is getting worse because our society permits employers to have a chilling effect on the online speech of their employees and those who might be their employees someday. And because the independent internet has shrunk dramatically, the handful of corporations that control almost the entire internet can now get away with siloing it off so that's what they're doing.
Discussions moved from independent forums to Facebook groups and Reddit. Discord later. Facebook and Reddit are not more anonymous than independent forums.
The decision to limit the API might therefore actually be a good for the average user if it stops archives such as ceddit and unreddit to show your posts after you've edit or removed them. Most people probably don't know about them though and believe their content is gone just because they click delete.
Facebook and Discord groups have been private since as long as I can remember. Sure, someone might leak if you invite the wrong people but compare it with Twitter where you can find all posts from user X including keyword Y. Or even do it through Google cache. It takes minutes to find something to use against anyone if you really want to. Based on the last election in my country, it can be something as simple as listening to the wrong artist when you're 14.
> our society permits employers to have a chilling effect on the online speech of their employees
More generally than that. Dylan Mulvaney made one beer advert while being trans and got death threats. She would be counted as a "contractor" rather than an employee.
The second they paid people to game their search engine (ie adsense), the entire thing was on the road to uselessness. There always would have been that incentive, but Google turbocharged it, with $33B of incentive in 2022.
From the point of view of a shortsighted company this might be a false dichotomy because if ad revenue from search grows, search is fine: neither rotting nor broken.
You are assuming that Google treats users as customers rather than as an exploitable natural resource.
Wikipedia has a downloadable data dump that would cost almost nothing to serve to Google, and it has an organisational mandate to make that data available. If they decide to charge for access to that I'm sure Google can afford it. Let's not throw around completely unrealistic hypotheticals.
How about asking for 10% of revenue when Wikipedia data is used to show search results? Would Google agree or starts another rival to wikipedia and kill it after 4 years?
Systems are so complex that, it is very difficult to predict how they would behave given parts of system gets impacted by other constraints (Reddit vs OpenAI data scraping, Reddit's urge to make money, Reddit moderators protesting against rules, 2 day blackout extending to infinite, Google releasing LLM papers, which are impacting its own business through ChatGPT and so on)
Despite the recent actions of Twitter and Reddit, last ditch attempts to maintain any relevance in the public consciousness, most organisations act relatively rationally at a high level, and Wikimedia has the added bonus of not needing to make a profit or being publicly traded.
Then my Puzzle Game Redactle (https://redactle.net) would be threatened much like what happened to GeoGusser being charged for Google Maps.
Related to the article: I've also had trouble ranking Redactle on Google because there are a few poorly implemented ad-ridden versions which are part of link farming groups. Many of them pretending to be my 'Redactle Unlimited' brand. Google loves a bunch of spam sites linking to each other more than thousands of links from authenticated users on Reddit, Facebook etc apparently.
Someone has to pay for these things to be developed and operated and typically when they do they get to call the shots, which leads to the modern UX disaster.
At least now the VC situation means we have a lot less product-dumping-by-any-other-name intended to destroy the market for legitimate participants.
At first, TV (broadcast) was free, with ads.
IIUC, when cable TV later rolled out, it was ad-free. Viewers paid for it instead.
But eventually even those paying for cable TV were shown ads.
So one lesson is that paying for a service might only delay its enshitification.
EDIT: It seems I was wrong about cable TV starting out ad-free [0]. So my example isn't strong evidence for my enshitification claim.
Streaming services. Midjourney. YouTube. Twitch. Even ChatGPT4.
I would gladly pay $10/mo for an information-centric search engine that filters out all blogspam and supports power queries.
I didn't know that LLM training sets used Wikipedia. I would have thought that the CC-BY-SA license (on all text) would keep them away. It's not like Chatbot-4000 can cite all the Wikipedia articles used to train it, nor is Chatbot-4000 going to license all its output under a CC-BY-SA license (as required by the license terms).
Are you suggesting that: "free knowledge, open source and open data, transparency, privacy, and collaboration" are ideas that "a lot of folks might not necessarily agree with?"
(I fully admit that I may be missing something from the few minutes of research)
I am interested in this because I believe that the users of social platforms are far better served by non-profits like Wikipedia instead of the enshitification that is fueled by the profit motive.
That's very far from reality. The money goes to Wikimedia Foundation, which is swimming in cash and reserves only a fraction of the money to run actual Wikipedia. That organization now has hundreds of highly paid administrators. Not a cent goes to any Wikipedia contributor nor does this staff directly contribute to Wikipedia itself at all.
Out of this Wikimedia Foundation, they created some spin-off organization that answers to nobody and gives away grants. To what a conservative commenter might call "woke" projects without any tangible return on investment.
But the point isn't whether you agree with that label or not, the point is that the organization is incredibly misleading in what it asks donations for, how their financial state is, how their organization really works, etc.
It's quite similar to Mozilla. You want to support Firefox development but instead the money goes to a foundation which spends it on lots of well-intended but generally useless projects.
It’s not a web search engine, it’s an ads search engine.
There was also a time you could paste a URL to a .GPX file into Google Maps and have it loaded. Quick and easy, well before OSM and Leaflet.
Way before 2020. Circa 2008 it began, and it simply continues to the present. ChillingEffect's downfall in 2015 was a massive blow to the public.
https://searchengineland.com/anti-censorship-database-chilli...
The other day I tried to google search 4chan, imgur and reddit came up before 4chan. heck 4chan wasnt even on the first page...
On a similar note, they are trying their best to hide wikipedia, now I have to specifically mention wiki most of the time.
I have the best website for a specific thing, I do well in quite a few SEO searches. However, the generic term, (which I'm still the best in), I am nowhere to be found. Blogspam and lower quality advice are all over the first 2 pages. Heck, there is actually terrible advice in the first 2 pages, like dis/misinformation.
The other day I was searching something I knew existed, I knew the page, and google would not give me the page. Ended up typing in the website.
Whatever is going on at google is in the major red flag territory. We need to perma fork android, we need to get off chromimum, find an alternative to Pixel. We are near the end times of google, they are going to be AOL taking advantage of old people and those who refuse to change their ways.
I firmly believe that many of the algorithm changes for choosing results are based upon ideology, and a desire to push commercial things down below the first page as a money grab forcing people to pay the ppc auction extortion.
Which means worse results for most searches beyond 'local shop open / phone number' - google, the new yellow pages.
New smaller, niche engines; less censored, more federated is what I'm looking forward to.
Can confirm. Same for me.
Just an asumption here. But could it be the fault of the web rather than Google's?
Could it be that the web itself is being so flooded with SEO-ed crap content that even Google can't sort through it all?
I'm starting to believe that some sort of Yahoo model is the best solution because you want a lot of source material to search through, but it needs to be context-aware so that searches are meaningful. I am vastly more effective in my job when search results return reference information rather than blog articles about somebody's weekend hobby project that only explores the most basic parts of a problem I'm working on. And this extends to other categories in my searching as well.
I'm also personally wondering if Google has realized (even implicitly) that because of all of this, it no longer has to do deep searching anymore and can save tons of money by regurgitating cached copies of "related" garbage because people will give up before any expensive operations take place. It makes me wonder if Google's internal resource accounting has anything to do with it but that's just pure speculation on my part.
Here's something to try: think of a really obscure TV show you watched when you were young, even better a specific quote. In my experience Google was pretty good at finding sources on where it came from. But now it's incredibly rare.
Disclaimer: I work at Google, but not on anything to do with search, and used DDG for my first year here out of habit. I have no inside info on search.
I don't find Google search to be bloated with ads. If I search for hotly competitive keywords then I do see ads, but the same is true where there is ad inventory on other sites, and when there isn't ad inventory that's not the true UX and it's only a matter of time until the ads arrive.
However, the majority of my searches are fact-finding, and I get facts, with no ads, or I don't even see the ads because I've got the answer before I get to them.
This is why I'd like to see a deep dive on this stuff, because the HN opinion is just not one that resonates with me at all at the moment.
Searching any programming terms, like camel-case variable names from a system library without exact quotes, just gives random results about Katy Perry and stuff now.
Granted, vectorization could be one of the issues plaguing search, but it isn’t the primary issue.
What would you say is the primary issue? Personally I see vectorization and similar fuzzy techniques as the things that have most affected search results in the last <10 years, both positively and negatively. On one hand it allows me to search really dumb queries and still get sort of relevant results, on the other it makes well thought query tuning for more specific results seem futile. I don't know what goes on internally, but it's as if Google just discards as noise most of the words in the search query. "Verbatim" mode seems to tune that down somewhat.
Because Google specifically has been degrading from self inflicted cuts for years.
Agreed that reddit going dark and machine generated "content" will only make it worse, but perhaps he should have talked about search generally then?
I've not seen very much discussion about the issues these LLMs are going to cause, this article is sort of talking about it - twitter and wikipedia could restrict their content, but - what about the millions of webmasters who ultimately created what the web is today?
What is the motivation of anyone to really create content any more, apart from for actual social discussions on social media? LLMs will steal the content, and no one will read it.
I really don't know what this will bring in the future. Will online content fizzle out in 2-3-5 years and fresh content will just disappear? Will LLM only be good up to 2025 data?
This is a massive problem that changes the web as we know it, yet I don't see any discussion of this. Mark Andreeson bought it up for a sentence on lex Friedman, but that was it. He didn't know the answer.
Making websites to sell stuff and services for example, ie a business. You'll find some of the best information on websites that are a business, because it's in their direct monetary interest to get you that information. As opposed to ad-supported websites.
How many of you remember back in 2005 or so, when the first lawsuit for delisting was brought against Google by a small business owner? The Google attorney showed up with a complete history of this person's searches and proceeded to demean him in court.
Probably not many of you, because Google buried the story. Today it can't be found (I can't anyway). It was written about in a highly critical book of Google a few years later.
Think Reddit's API story is bad? Google did the same thing back in 2010-2012, slowly tightening the clamps on anyone monitoring them to put them out of business. My company was saved by a call from a famous billionaire to Matt Cuts, after which our little problem promptly disappeared.
The reason I remember this is because I wrote three books about Google, and I was sorely tempted to include the story but my publisher decided it would wiser to not include it.
Has there been any actual analysis showing that Google Search is getting worse nowadays? All I'm seeing are anecdotal reports and feelings, and your link shows that this sentiment has been around for more than a decade.
It actually impressive that google has managed to avoid this flaw for so long.
The only thing I am not sure of is have advertisers finally found ways to game the system or is google doing it to themselves having forgotten why we all started using it in the first place.
I reckon that chat GPT will end up like the Douglas Adams HHGTG dystopian elevators at some point too :)
The author is overlooking that Google can cut a deal with Twitter and Reddit and everybody else. That's for mutual benefit. It keeps driving traffic to them which is great PR and it keeps Google relevant.
I also use it, but if Bing starts to make real gains and Microsoft thinks more DDG users will go to Bing than Google if they kill DDG, they will kill DDG.
They’ve been playing a very risky game for a long time.
In the old days, we ran our own personal websites and we want our content to be discovered, in contrast with today where someone else host our content and wants to have a lot more say on what is discoverable. I think we need to have less reliance on centralized hosts if we want more control over searchability of content.
This is going to be the interesting part to me - one HUGE usage scenario for AI would be doing automated searches and distilling results from the web - if sites start preventing that then AI innovation could be stifled - OR we'll see a scenario where the average person is very limited with that their AI subscription can access. We could have a situation where e.g. Fandango prevents AI from searching its site to help people plan their movie outing, but instead has their own model that they'll charge for access. We could have models talking to each other, deals made between the owners of the models for access, and the average citizen uses a search engine optimized for monetization, that may have data that's months out of date but they can pay extra for the model that provides current data.
Before, there was a very strong if not existential incentive to get indexed by scrapers, in particular search engines. Now the situation is flipped where there's a strong incentive to put up ever higher walls, in particular for LLM scrapers.
The difference is obvious: LLMs don't give back. They just rob your life's work and monetize it locally. Why would a platform or its users open up to that?
Wikipedia does charge Google.
https://timesofindia.indiatimes.com/business/international-b...
So now it is very rare to get the answer on the search page, because it makes no money for google.
I've seen the other side of the coin, far too frequently. Many websites allow Googlebot to index for free, but once a person clicks on Google serach results, they ask for money.
SEO becomes a lot more difficult to pull off when there isn’t one search algorithm used by 99% of users.
Which means we return to late '90s and early 2000s situation when there were no behemoths hoarding data and knowledge and there were many specialized websites. I am happy with checking tech news on HN, checking stuff about cars on a car website, read about photography on a photography website and so on, instead of searching countless Facebook groups and subreddits.
Lots of smart people making very bad products that get in the way of other smart people making good products. It’s shocking to look at the wreckage of their decades of poor decision-making.
The idea of sites making their data inaccessible is not new. Experts Exchange is still out there siloing the answers but Stack Overflow has all the market share though being open. Google will be fine.
Is google's search business dying? (thepoc.net) 2010 https://news.ycombinator.com/item?id=1217737
Is Google slowly dying? 2014 https://news.ycombinator.com/item?id=7779020
Is Google Dying a Slow Death? 2010 (http://brooksreview.net/2010/11/death-google/) https://news.ycombinator.com/item?id=1929821
It would be good not only for the customers, but also for the ISP since it has less junk to travel through its network.
Why is every search result for even the simplest yes/no question a 1500 word article with 2-3 picture and an embedded video? Because that's what Google thinks is a good search result!
Google penalises your search ranking if users quickly navigate away from your page - this means if you have a page that quickly an efficiently answers a searchers question you get penalised whilst the site that buries the answer two thirds of the way through pure blog spam gets rewarded.
Easy, rehost it.
For content creators this has meant they need to find more and most importantly more agressive ways to monetize their content since less and less traffic means you need to show more ads to keep the money flowing. If google lowered the amount of answers it actually gives and would lead people to click on the sites more there would be less ads and users would be happier as well and there would be less paywalls as well.