Google Memory Loss
tbray.org
tbray.org
To add insult to injury, if you do try to make complex and slightly varying queries and exhaust its result pages in an effort to find something you know exists, very often it will think you're a robot and present you with a CAPTCHA, or just ban you completely (solving the CAPTCHA just gives you another, and no matter how many you solve it keeps refusing to search; but they probably benefit from all the AI help you just gave them, what bastards...) for a few hours.
Google had the biggest most comprehensive index for many years, which is why it was my sole search engine. Now I'm often finding better results with Bing, DuckDuckGo, Yahoo, and even Yandex, but part of me is very worried that large and extremely valuable parts of the Web are, despite still being accessible, simply "falling off the radar".
I think the biggest irony is that the web allows for more adoption of long-tail movements than ever before, and Google has gotten significantly worse at turning these up. I assume this has something to do with the fact that information from the long tail is substantially less searched for than stuff within the normal bounds.
This is a nightmare if you have any hobbies that share a common phrase with a vastly more popular hobby, and is especially common when it comes to tech-related activities. I use Linux at home, and I program VBA at work. At home Linux is crossed out of most of the first few pages, and I just get a ton of results about Windows, and at work VBA is crossed off and I get results about VB6 and .NET.
Completely. Useless.
I can only imagine this has something to do with their increasing reliance on AI, and the fact that the AI is probably incentivized to give a correct response to as many people'above the fold' as is possible. If 95% of people are served by dropping the specifically-chosen search term, then the AI probably thinks it's doing a great job.
It seems like the web is being optimized for casual users, and using the internet is no longer as skill you can improve to create a path towards a more meaningful web experience.
"It never rains on a Wednesday in Rockshire"
..and Google would return websites containing that exact phrase. That no longer works. Nowadays parentheses are largely ignored as far as I can tell. This is super annoying because quite often I am really looking for a website/websites containing a specific phrase.It seems they internally switched to indexing single words only. So if you search for a phrase Google will instead return websites containing (some of) the words in your phrase in no specific order and maybe not even next to each other.
I think I understand why: Indexing / searching based on words is massively easier and massively less resource intense than a system which can search for specific phrases.
Furthermore the type of searches Google wants/expects you to do e.g. "best hairdryer" work well enough, in fact work better if you only search for single words and then filter / organize the results using AI / using information you have collected about the user.
EDIT: I was wrong. It still does work for many phrases, just not the ones I tend to search for. See below.
EDIT 2: I actually have no idea what is going on there. I know that searching for exact phrases doesn't work for me like it used to but I have no idea why..
Lately I've also been noticing that the behaviour of some google search operators are broken. " something " "otherthing" is not considered as an AND " something " OR "otherthing" is not considered as an OR. Google shows me the results it wants. I recently tried to research FreeBSD and Meltdown (I tried many times: "FreeBSD" "Meltdown", "FreeBSD" AND "Meltdown" etc.) and almost no result involved the terms both FreeBSD and Meltdown. The interesting thing was, Google did not say there was no results matching my criteria, it kept showing me popular IT news, linux news etc. It was extremely frustrating.
The only operators that work are site, date and filetype. The logical operators do not work reliably or do not work at all.
If you're searching for something popular google finds it best, but if your search patterns are deviant google ignores you and even thinks you're a bot and refuses to service you.
Today, if I'm searching for something unpopular or specific, I usually get frustrated. You would expect the opposite to happen as the size of the web should increase over time.
This has been bugging me too the last couple of years. Sometimes I would actually prefer an empty answer to one where various words are missing. Empty can mean an idea is still unexplored.
I've started sending feedback to Google every time this happen. Maybe it'll make a difference if more people do it?
What’s particularly infuriating is when searching with verbatim still ignores keywords.
They even have a patent on that: https://www.google.com/patents/US9407661
I've been lazy and so acclimatized to the UI that I haven't changed, but I plan to now. It's especially difficult finding medical information. Mind boggling how 6-7 years ago it used to be so much better.
Google by itself isn't very good, but Googling from the Google Chrome address bar these days is something else.
I have the same problem.
To pass Google CAPTCHA you have to perform like an average human. This is different from getting it right. Once I started being lazier, (e.g. back of a sign isn't a sign, picture that contains a storefront but in the distance doesn't contain a storefront), I have had greater success rates.
It's like Google is training me to be dumber.
I'm not certain yet, but I think it has to do with systems relying more and more on the the system giving a high weighting on popular or frequent searches. In other words, these systems are filtering content by how frequent they are searched and annealing returns which have low frequency.
This makes sense from a machine learning perspective. If I want to build a system which returns a search quickly, then it will be biased towards pathways with strong weights. So in the end, systems would be biased against outlier searches and very specific terms which have low or weak search pathways. Effectively the system is getting really fast and good at giving you the most popular return, rather than a precise return.
The major problem there is that over time the search space will atrophy, much like memories do, and will kill off pathways which have low frequency. It's unclear if this is good or bad long term for a search engine, because it will remain popular for the majority of users, so long as their search terms and desired results live within the same space. In other words, we're creating a less diverse and more homogenous space by virtue of giving higher weights to more common thoughts/searches/desires.
I've always double quoted the specific term I'm looking for. That usually bypasses this, ie. searching for "foo" will look specifically for foo and not food or what have you.
Coincidence? No...
This is the algorithm deciding that it’s way results in more ad revenue for Google. What do you do? You repeat the query again and they show you a second lot of ads.
But let’s be honest if Google could show you as many ads without returning any search results at all, they’d do that in a heartbeat.
I haven't done this methodically, and I can't prove that this is happening, but it's infuriating nonetheless.
I even have four photos I took at night, in burst mode, as a Bughatti Veyron zoomed passed, and yes, it can recognise those...
That's why I'm back on Firefox after quantum release. I hope mozilla never, ever, ever does something like this but I remember seeing something similar on nightly once. It gave you search suggestions first, that redirected you to google, with option to disable it in settings.
Another really interesting thing I've noticed in gmail relating to search is that the number of matches for a given search is approximate, which makes perfect sense if they're using some kind of probabilistic data structure. However, when the correct number of matching emails does become known, because you have gone to the end, the result is not cached even client side. This gives a weird effect when combined with pagination: you go back a page, and the number of matches changes to the estimate again despite the fact the actual number is now known.
Seriously, this is egregious. You rely on your email provider to accurately search your inbox - some emails are important business, tax, and legal documents that are relevant for years, even decades. Or at least be fucking transparent about the fact that you are not really searching all emails. I know Gmail is a free service and in the T&C you agreed to (figuratively) sell your soul but this has huge real-life implications.
I don't mean to pick on you but picturing the perspective behind this comment is very funny and a little sad to me.
grep is almost 40 years old. It is free software, fast, and doesn't share your data with anyone. Small knowledge of the file structure of MIME enables more advanced search. This is all without mentioning desktop-based email clients.
Reading your comment, I can only picture some web-page javascript-based track-you-and-show-ads 15-employee company whom you give your email password so they can connect to another service and make high-latency queries on your behalf.
Shows how far we've come?
Just exclude Mail from Spotlight search in Settings.
Well... this is why you might want it. Your data under your control. Your choice of tools.
If you're using GMail and Google decides to turn GMail to crap, well, bad luck.
I'm talking about OWA Search in O365. I use mostly VDI these days, so PSTs are out. It's a frustrating issue to me because OWA search is better in many ways for more recent stuff.
Also note this is an anecdotal interpretation based on my experience.
Almost there...
But now we are at peak AI hype cycle. No wonder it's gone downhill. I'm sure the AI does better on whatever metrics they tell it to maximize. But AIs game the hell out of metrics. And we are still nowhere near human level intelligence. It doesn't understand your query or the content of the websites. All it has to go by is simple keyword matching and meta indicators like the size of the website.
So now basically all searches return the same handful of large websites. When was the last time you got a search that went to some niche forum? Or some small little homemade website by someone passionate about that specific subject? No, it's always a wikipedia link, followed by a bunch of contentless news articles. And there's never any point of going past page 1, because every other page is like that too.
And now the web has become this: https://www.ncta.com/sites/default/files/platform-images/wp-...
I think this has dumbed down society because it certainly has dumbed down me. Also, complicated topics are dumbed down to 1 sentence answers and it's very difficult to get detailed information about something. Ironically, I've started to go back to purchasing and reading books if I really want to go in depth on a subject.
I wanted to see how Raspberry Pi assembly worked - quite a niche subject (who programs? In assembly? On a Raspberry Pi?).
Yet I found quite a few blogs going through the subject - a tutorial of sorts.
I was also looking up how to write an OS (just for fun) - another fairly niche field. Yet I found tons of tutorials in C, C++, and Rust.
I remember trying the same in the 1990s.
Nada.
You couldn't find nothing. I mean maybe you'll find a basic site with the source-dump of an OS (often without building instructions), but you want to dig into the meat (here is how you get from bootloader to x86-64 in 21 days, and why it works)? Foggetaboutit.
I definitely don't want to go back to pre-google web1.0.
I suppose that this is a use case where ML could help a lot to recognize "content farmed + SEO optimized" content. Hopefully Google could improve the situation in the future - supposing they are trying.
The annoying thing was that the content farm appeared _above_ the original content.
There was actually a result for a small wiki of pi enthusiasts. But it was buried 10 links down on page 2, and mixed in with all the other noise. There's also a lot of news sites and blogspam (a lot of links have a little indicator that says "x hours ago". As if good articles expire.)
Both Yahoo and AltaVista were started in 95.
A major issue, that affected finding "Raspberry Pi assembly" in the 1990s was that the Raspberry Pi didn't exist yet.
A lesser issue, and perhaps one that may have made it difficult to figure out how to "write an OS (just for fun)" is that few people had done this and published tutorial-style material for absolute beginners on the web- and most of those tutorials that did exist weren't on the world wide web (but on e.g. usenet).
To be honest, I kind of like that part. Most of the time, if there are no useful links on the first page, I know to refine my search terms right away, instead of wasting my time clicking through pages and pages of useless stuff.
I remember those days too. It did feel like Google search skill was a super power. I also remember my non technical friends and family being pretty much unable to find what they were looking for online. Google did not turn their rambling approximate queries into the results they want.
Now I find that although the precision of the old style is gone, Google is incredible at guessing what you want. Back in the day I might find the page I was looking for on the twelfth 'o' of 'Goooooooogle'. Now I rarely have to venture outside the top 3 results. I find that my family no longer needs my help to fast craft a search query.
Isn't it possible that Google just made the choices that are better for the average user and left some of us advanced users out? If your response to that is "they could leave me an advanced mode", consider how much work it would be to maintain to serve a tiny customer subset.
I think if you want a search engine for power users you want DDG
Disclaimer: I work at Google but not on Search or anything close.
More like guessing what most people want.
If you want something specific, sometimes you re out of luck.
I definitely miss old days of Google. After Amit Singhal left, things haven't been same at all. He had resisted unexplainable AI getting in to search features. But as he left, RankBrain had Ai-driven feature that is 3rd most significant. That feature is the reason why you often see pages even if they don't contain keywords and even a phrase you had specified. Those old guard knew that trading explainibility and little bit of revenue with slight decrease in customer sat wasn't worth it.
It takes a different design of index and algorithms to do what Google has openly had as a goal since well before Porat came on board than it does to be the kind of highly literal search engine Google started as.
There's a market for both, and the market for a well-designed, comprehensive, highly literal engine probably pays a lot more per user than the kind Google is focussed on. But it also is much narrower if there is a Google around. (It's also not clear that the web is the highest-value corpus for such an engine; the really commercially successful ones are more specialized and have large, human curated and annotated datasets is specialized domains, notably law.)
Whenever I see horrible search results, I often wonder if those people who work at Google, and surely use their own search engine as much as anyone else, have noticed the degradation and what they think of it --- especially those who are in charge of or even working on the search engine themselves. Does Sundar search, get horrible results, and think "Why is my search engine half-broken? This is my company's most prominent product, and it's not working as well as it used to." No doubt it affects all their engineers too, the ones who will tend to be looking for the most obscure things.
I think your point about metrics and revenue is very true --- they are numbers that can be easily compared, while the quality of search results is not (and also subject to a lot of different conditions); being a "data driven" company, they obviously place much emphasis on the former, ignoring the negative but not easily quantifiable effect on search quality.
As the catchy saying goes, "Not everything that counts can be counted, and not everything that can be counted, counts." Unfortunately a lot of Google's management don't seem to believe in it.
"I had been playing the accordion Davy lent to
Rosie during winter break"
"The language they're using is not that different
from the one I wrote PlayGUI to use"
"I've been playing a decent amount of music lately,
mostly guitar and piano."
"warm dry socks was the most important aspect of the
festival"
"This wouldn't be that bad, if it was not exactly what
happened a year and a half ago."
Google found all three. For each one there were either 2 or 3 results: first my old post, then one or two from rssing.com which seems to do something with my rss feed.Trying them with Bing, it also found all five of my posts, and ranked them first in four cases. In the fifth case ("This wouldn't be that bad, if it was not exactly what happened a year and a half ago.") it ranked a goodhousekeeping.com post higher, which had all the individual words but none of the phrases.
(Disclosure: I work for Google, though not in search.)
In fairness to Google, that's always been a problem, and even in the good-old-days in which keyword-based searches were more effective, there were content aggregators that would copy the entire contents of phpBB-style bulletin boards (and USENET newsgroups) in order to rehost them and get clicks.
On the one hand, I want to say that it's precisely the sort of SEO/spammy practice that Google should be deprioritizing in search results. On the other hand, sometimes these copies/mirrors of content are the only extant copies of content when an original blog goes away. Although the motivations of the owners of these sorts of sites may not be as pure as that of archive.org, the result for the searcher is equivalent: the desired information is found even if it's only a rehosted copy.
All five? Or 3/5?
Sometimes the datacenter your search query lands in might not have a copy of the necessary page. Now they have to decide if they delay the entire search query to remotely query another datacenter, or not. I would guess returning the results early is nearly always more important than returning a result which is so rare it has never been clicked in the past decade.
However, what I am getting at with that simple example is for the searches to be quick google keeps these lists small. So there is a limited space due to time-constraints. So google must decide what is relevant for the available portion of their index.
However, that does not explain why other search engines don't have trouble with older sites/links. I suspect it's more of business decision than a technical one.
edit: It's not tedious for me on my browser, just click Verbatim on the LHS of page. (Can select that or All results)
"lkasdfjer" + "samsung galaxy s8" : brings up this discussion
That's not relevant to the article, which says that the results are not available AT ALL. (Although as of my posting the two articles seem to be available again.)
Google Search no longer runs a clearly defined algorithm to find search results. It is a collection of AI systems that are trained continuously on a variety of data. There is probably no human alive who fully understands how Google makes decisions about which results to return and how to rank them. They just understand how to provide feedback to adjust results they don't like.
If the system gets rewarded for finding common things quickly, then it will adjust its internal algorithms to make that happen--perhaps even if that means dropping unpopular results altogether.
It would be like if our society couldn’t fully function without Roman concrete, but no one knew how to make it anymore. Which is sort of what happened during medieval times: people continued to use Roman roads and Roman bridges. But if anything degraded, there was no hope of repairing it.
Today, it doesn't matter how you build your query, Google returns good results in any case. That also means that you can't search for specific info by phrasing queries differently but for the vast majority of people it makes life much easier.
Even Google falls into the trap that a worse solution with more fashionable tech gets deployed.
Not in those words, but they do claim to aspire to “Organize the world’s information and make it universally accessible and useful.”[1] which ought to include old web pages. They've gone to the effort of finding out of print books and digitizing them to make those searchable so it doesn't seem like a ten year old web page should be such a stretch.
# robots.txt web.archive.org 2013-10-02
User-agent: *
Disallow: /
User-agent: ia_archiver
Allow: /I don't get that line of thought, somehow people have starting defend lack of quality as something expected or reasonable.
The whole point of going to google is to find stuff, that includes "boring" old and obscure stuff that won't sell ads. But that is part of the deal, if google don't care about that why should I care about google?
While we are on the subject, I still miss being able add a + in front of a word to highlight its significance, this was removed in favor of google+ (you can still do it with quotes) but now when google+ has been irrelevant for the better part of a decade maybe it is time let us quickly emphasis words? Sigh.
Maybe a competitor will come in and steal marketshare from Google by filling the hole Google leaves behind when they increasingly make changes which annoy a subset of users (duckduckgo is probably closest, and I use it as my default search engine, but I don't find it's as good as Google yet in a lot of cases). Maybe enough people will switch to be an issue for Google, in which case Google miscalculated the cost of pissing of that subset of its userbase. Maybe so few people will switch that it's offset by the cost savings of not keeping everything indexed, in which case it might've been the correct decision, from a capitalistic point of view.
This focus on constant growth is the issue with relying on companies and capitalism, but that's a bigger discussion.
This hasn't anything to do with capitalism. It is pure greed. Companies willingly do anything for a slight increase in revenue even if they willingly acknowledge that it will cost them ten times as much in the (not so) long term.
There is nothing about capitalism that says you must be a colossal idiot, that's just a consequence of a poisonous culture where employees don't give a crap about the company but only focus on their own career.
If my search bar DDG search doesn’t return satisfactory results, I hit my address/search bar key (cmd+l for me), press left arrow/start-of-line key, type “!g “ and hit return. That gives me a google result page quicker than the time it would take to reach the mouse. It’s become a good workflow.
I’m having to use it less and less. DuckDuckGo is getting really good.
Google seems to have decided that Wikipedia is the only blessed noncommercial source of intelligence.
I guess, if I were to put it strongly, I'd say: using Google is not like using the Internet any longer.
FWIW, HNers may wish to check out Yewno[1], a knowledge search engine based in Redwood City that I've had the pleasure of being (tangentially) involved in.
[1] http://yewno.com
[2] Yes, I know this is indexed. It just frequently gets buried in my searches.
The site's robots.txt seems permissive enough. What's up?
(I agree looking for one's own articles is a specific case - in most other situations you'd want to know Google's reasons very badly.)
A good review would show more profs than just websites behind the same one domain.
Interesting though, let´s see if there´s more news on this.
Uh, five minutes after I first tried, now the review is the first hit for [lou reed "rock n roll animal" tim bray] and a bunch of variations. Enough people searching for it might have changed the state of the system.
A: I don't remember.
Consistency, availability, tolerance: pick two (https://landing.google.com/sre/book/chapters/managing-critic...). If you think about the uptime constraints of www.google.com, you'll possibly conclude that the limits of distributed systems necessitate a solution where some data is transiently unavailable.
The same content hosted on another domain would appear much lower on Google than a Pinterest post (or a post on any other large website). Not sure if that's really the best approach but I guess it's not easy coming up with a better one.
Google is not a search business despite opinions to the contrary. A search business would have as its main source of revenue customers paying for either search results or the search technology itself. Google makes its money via the delivery of ads, through its own advertising sales and placement platforms and through the audience it can provide from its own content.
Google _seems_ to be returning more popular/mainstream sites at the expense of less popular sites that may be more relevant.
Also- Google has stopped penalising sites that don't contain all of the search terms. It has always removed "stopwords" (words that are too common to be relevant: and, it, or, etc.) which is fine, but now it seems to remove significant terms as well. This makes a big difference if you are searching for programming stuff, particularly error messages.
From a pure search point of view Google is losing ground to its competitors, and that has _never_ happened before.
This shows the value of actually grabbing content that you plan to use or hope to refer to in the future, rather than merely bookmarking it. And it also underscores the value of the Internet Archive.
Yes. Everyone should install the Wayback Machine plugin and click the "save page now" whenever they find something useful or interesting:
Chrome: https://chrome.google.com/webstore/detail/wayback-machine/fp...
Firefox: https://addons.mozilla.org/en-US/firefox/addon/wayback-machi...
I hate hitting an unarchived dead-ends when doing research, so I'm trying to do my part to prevent it. Many page I've archived had never been archived before I saved them.
https://encrypted.google.com/search?hl=fr&q=%22Dans%20mon%20...
https://encrypted.google.com/search?hl=fr&q=%22Eh%20bien%2C%...
BUT, interestingly enough, it can find posts for the very first months of the blog existence, but stays clueless for several posts dated a few years later. For instance, several exact strings from this page are not found by Google :
In my case, I would like to use Google search to bootstrap a few minor search engine indexes and collect data for NLP projects. But the free version is too limited and the paid version prohibitively expensive, so, no luck.
Shoot me an email, I’ll hook you up with 1,000 credits. It’s funny one of my first use was from ML as well, but it end up being used mostly for SEO and web marketing.
I'm sure you could start client-side caching every search you ever do to Google.
But if you're searching enough to eat up Google's bandwidth, they're paying for that data and they're under no obligation to keep serving you as a client (much as any server is under no particular obligation to serve a search spider).
Do you not see the irony?
Search engines don't allow people to scrape them (resorting to blocking after scrapers ignore robots.txt) because they don't get anything similarly valuable in return.
(Disclosure: I work for Google, though not on search.)
Search engines crawling millions of sites each with---on average---a few MB of data distributes cost globally.
Extracting terabytes of index data from a single search engine's repository consolidates the cost on the back of that repository's bandwidth provision.
These are not symmetrical cost structures.
Usually GOOG had best results for technical queries --but fresh results for those tend to be the better results (things go out of date pretty quick because of quicker release cycles)
However, on occasion it has been difficult to find very specific results to non technical questions --I never even bothered with DDG or Bing, but now I will surely give them a try.
If they are data driven (and they are) for their average users, long tail, old results probably don't make sense --how many people really go down to page 20 of the SERP? I'm sure it's a very miniscule number.
Unfortunately those people are the ones who are searching the hardest for the most difficult-to-find things, and thus need the services of a search engine the most. It's unfortunate because, for every one of those, there's probably millions of others who just want to search "facebook" and click the first link; a site that I don't even use yet can recite the domain name of off the top of my head.
This is a contrived example, but say you want to look something up regarding what someone said about drug overdoses back in the 80s. Google would try to insist on brining up information about the most recent overdose studies for example, because people are currently discussing that more, so to Google, obviously I should also be looking the fresher content, so the end result is I get less relevant to virtually irrelevant results.
(word OR word) AND (word NEAR word)
and get excellent results. Of course, the Web was much smaller then.
And that's kind of the point, right? Beyond a certain threshold of popularity, some things aren't always available from every search query, because a distributed system can't have 100% uptime, consistency, and tolerance to network partitioning.
It's well known that there are a bunch of metrics that go into which results to return — metrics including things like pagerank, (probably) historical value (# of clicks when the page appears in results), and social media popularity.
I wouldn't be surprised if Google has experimented with training models to predict most of those metrics, given only content from the site itself, and tried using those models as a filter for what to index in the first place. If the NN is accurate enough, they can use it as a filter at indexing stage ("should I index this?") rather than at the results ranking stage (where real data, rather than NN model output, answers the question "should I show this page close enough to the top of results that someone will see it?").
And that's where I don't agree with the author. I've been thinking about this and to me it seems we need to try to foster a culture of forgetting. Just because we can store everything forever, doesn't mean we should. That kind of thinking is exactly where this 'track everything everyone does'-mentality comes from, governments seem so keen to apply in the name of 'terror prevention'. It also values regular stuff you do way too highly.
Let's face it: A huge portion of the web is garbage and most information on it is ephemeral. The really useful stuff, like encylopedias, will be used regularly and thus keep indexed anyway. For the rest: Just let it go. Forgetting can also be relieving, you know?
* Is there a metric of search quality which is appropriate here -- specifically, "when I search for [site:tbray.org rock roll], and receive a set of results, that set includes Tim's article"? What do we call this metric? The metric would be lower when the result set is empty (no relevant results returned) and higher when the result set contains the desired article (a relevant result was returned).
* How would you assess the quality of this particular search against a metric?
* How would you measure the overall quality of "all searches in the past hour, including the [site:tbray.org rock roll] search"? How would this one failure to find a page contribute to an overall success rate?
* Is there any possible automation that would notice whether Tim's article has started to be missing from indexes and say "hey, this represents a loss of a kind of quality"?
* Suppose the index were to (say) discard all pages created before 1999 but simultaneously improve the relevance of all queries that find more recent results. If (say) 99.99% of queries have users happy getting only post-1999 links and (say) only 0.01% are unhappy because they specifically wanted a pre-1999 result, but things get way way better for the 99.99%, was that a bad change? would any metrics show a problem?
I don't see super satisfying answers to this at e.g. https://www.quora.com/How-does-Google-measure-the-quality-of... or https://www.quora.com/How-can-search-quality-be-measured . If I'm reading right, it sounds like part of the state of the art for search quality recently involved human raters manually running sample queries… That seems kinda crazy / totally unlikely to catch certain obscure issues. But then again:
* What is the service level objective for search quality? If search is getting way better for 99.99% of users because of various optimizations, is it a problem if a particular 0.01% of queries such as Tim's old review query, which he expected to find one specific page, instead find no results at all?
And then I guess I wonder:
* According to whatever metric correctly captures Tim's review being missing as a problem, what is the current search quality of Google web searches and how has it been changing over time?
- recall: number of relevant documents retrieved / number of relevant documents
- precision: number of relevant documents in result set / number of documents in result set
Given a query like [site:tbray.org "rock n roll animal"], and knowing that the 1 relevant document we actually want is the review at https://www.tbray.org/ongoing/When/200x/2006/03/13/Rock-n-Ro... , I think we can say that
* if Google search returns 4 results for the query, not including the review: precision is 0/4, recall is 0/1 (so p=0, r=0)
* if Google search returns 5 results for that query, including the review: precision is 1/5, recall is 1/1 (so p=0.2, r=1)
But while I _kind of_ understand how we can use these measures to assess the outcome of a single query, I'm really not sure I understand what meaningful ways are available to aggregate those metrics. Suppose we're going to get 1M queries in the next hour. Do we prefer an algorithm which has the highest mean F-score per query? highest median F-score per question? or which has the highest 1st percentile F-score per question (99% of queries get the best possible outcomes?)
If there is published literature on how search quality is measured I'd love to see it. Would be especially interesting to see real-time data -- e.g. what is the impact of 1 data shard outage on overall user-experienced quality according to some metric?
I'm not sure though how well it's kept up in terms of aspects like real-time search and graph search, both of which are fairly recent developments.
If I google "lou reed perfect night review", I stopped looking after page 17 of results. There's just too many results.
If I google /"lou reed" "perfect night" review/ with quotes as you see them, the review I published is on page 2, result #3.
I feel your pain, but, as someone who started publishing content in 1995, I don't see Google as having a memory problems.
My pain points are related to how Google crawls my content, my sitemaps, and how the two seem completely independent of each other.
I wonder if they have any other index snapshots stashed away somewhere, I would love to do that again. Even if only to retrieve the urls of old homestead websites I had back then.
EDIT: Yeah, the TLS handshake takes 20 seconds for me. I don't know why. Everything else works fine.
Information nowadays isn't dead like it used to be, it's alive. It have desires and seeks readership, it mutates and wants to spread. This information has its source in human intelligence.
It seems possible it is outwitting Google's AI as the smart people have shifted focus from being smarter than "bad" information to building dumber than human AIs to improve bean metrics and bask in glory.
The core problem has mutated from "what is relevant" to "what is quality". Pagerank and the web of dead info could answer both with one number because the quality signal and the relevance signal were the same and hard to fake.
But if you can hyper optimize for relevance by making content addictive thus affecting downstream attitudes on "relevance", quality is no longer relevant. Current AI is smart, but still childlike and easily corrupted/manipulated. It's a black box that can't be inspected and adapts to change but because it's tempo of dynamic modeling is slower than a real humans it can be "trained" or hypnotized.
Hell, it's probably the core sociopath skill. Being able to manipulate the value/principles system of someone/something else.
There are indeed many keyword searches in which google is obviously censoring the internet in big chunks. Because they are so opaque no one knows if it's directed by a gov agency, or shareholders, or the vision of some guy at the top who thinks they need to grow up, or some small team that is so hellbent on destroying some things they do not worry about the collateral damage - or maybe it is "machine learning" figuring out if you censor big chunks here and there that people will spend more on ads. Who knows? Very few people, and who is affected? Many people, many who do not even know it. Facebook of all systems is changing thier over algo ways to be better for people and less for the algorythm - will things change with big G now that the last babysitter left?
Google's current path imho is becoming the next yellow pages. Sure they will do fine being embedded on so many mobile devices and being the go to place for crowd sourced directions and locations - but the yellow pages was king for a few days and then people realized it's value to the consumer and to the business was no longer what it was and now who is using yp? More details on a non-google run blog soon. Who can I trust not to give up ip info of posters to tech peeps at G? Wordpress.com? Oh wait, they use google fonts and all kinds of things don't they. Hopefully people can help put this info together with me, it is prime time to create some new alternatives that only do slices of what G used to do well.
Furthermore, when someone links me to an amp.whatever page, I might be in an even worse mood if I have to talk with them about it afterwards. imo, Google and facebook algorithms are half of what's causing most of the conflict and hate on the internet these days. The machines have literally taken over, and they've started with our minds. YouTube suggestions have devolved into nothing more than a rabbit hole that gets scarier the further you follow. Either they are completely out of touch with what people want or the mentality of society/humanity is way more fucked off than I'd ever imagined it could be. Until they provide actually useful and configurable settings for people who can think for themselves, I will not return to using these products.
Besides, how fair would it be for me to freeload off an ad company's services when my hosts file hasn't allowed an ad to load in my browser as long as I can remember?
For those who don't do the search. (It's officially misspelled and instead referring Robert Riger, an art photographer, (hence the one leg crew :) ))
I personally don't feel that Google search quality has degraded very much if at all. It's true that I rarely get new sites outsite the echo chamber or few 1000 popular sites, but to me 99% of the time they give me relevant and useful information I am looking for. To be honest Google search results on average is still heads and shoulders ahead of the competition. in almost all aspects.
Another issue is that I don't know what the filter is selecting either, a 6 year old article might be better if it's been updated, but I can't tell from the interface what property is being filtered.
What drove me nut's, is that their QA didn't catch the bug with the date format order for years. It only recently got fixed. The calendar selection was regional and put in the date in the regions format (Often dd/mm/yyyy), while the query form expects mm/dd/yyyy.
It's not purely a function of 'date published' it is also about frequency of access
This behavior makes sense for most people, those who don't know what they're doing, at the expense of rendering the tool useless in cases where advanced functionality is needed. Perhaps Google should have an advanced, AI lite version.
Obviously, indexing the whole Web is crushingly expensive, and getting more so every day
Why? CPU, Memory, SSD has all gotten a lot faster and bigger. Bandwidth is also a lot cheaper for Google since they essentially owns the fibre. Faster Algorithm may have an even bigger impact.
I would have thought information, in terms of text ( Not Video and Pictures which is a few order of magnitude larger ) would now be cheaper then the old days.
By _improving_ their system, it creates some difference, if no one cares or no one can justify the necessity, then google doesn’t need to care.
The constant increasing amount of information on the web nowadays is certainly a burden to Google. And if such AI can handle 99.99% of what’s matter to user, this AI is already a brilliant one.
After all I don’t think google ever wanted to be an archive searcher.
My mental model of the Web is a human brain with all kinds of weird mechanisms - one that forgets things, makes things up, mixes memories. Google is like an "association cortex" - take it away and all the memories are still there but access to them is much harder and needs to go through more indirections (links).
https://github.com/synchrony/smsn/wiki
We can even interconnect them -- selecting what's private, what's public, what's shared with some people but not others.
If this about finding older parts of the web, don't forget to use the Wayback Machine of archive.org:
E.g "from:torvalds and to:linux-ext4" to bring up all emails ever with those properties. Add some free text and/or "tag:foo" to narrow it down.
I can find the article this person is referring just by looking at a string of it. So the article _is_ indexed. Not sure what he is referring to.
"Nice try!" :D
1) Bing is much faster to load
2) Bing doesn't muck up the links. I can right click->copy and get the actual link. With google you get a long incantation that doesn't tell you what site it goes to.
3) Bing handles a few things better, eg, "$500 in ₹" works in bing but not in google.
However, when I need local results, or results about something ongoing, Google is the undisputed king. Nobody else even comes close.
As far as I understand it (I never actually looked it up to confirm), currency codes = two letter country codes + first letter of the noun from the currency name, so they're relatively easy to guess right.
Google is in big trouble.
The hardest to replace part is search today.
Homepage of Joe Infinity
You are on page <?$pageNr?>
<a href="<?$pageNr+1?>">Next Page</a>
can not be completely indexed. A search engine will crawl it to some depth based on many factors. Age might be one of them. There is no way to index 'everything' on the web.The first sentence is just common sense, and no particular proof is needed. The last sentence might or might not be true, but the anecdotes in this article say nothing about whether or not its true. The problem is that we don't know how Tim selected these two particular pages as examples.
If he randomly selected two 10 year old pages from the universe of all such pages, it'd at least be a valid methodology, just with far too small a sample size. But obviously he didn't do that. If the methodology instead was to search for pages on Google first, then on Bing iff there was no Google match, this tells us nothing at all. You need to run all queries on both engines, not just the ones that fail on one search engine.
Another reasonable method would be to look at aggregate referer trends; is traffic from Google to old pages decreasing faster than traffic from Bing to those pages.
How is this common sense?
> The problem is that we don't know how Tim selected these two particular pages as examples.
Yes we do know. He was using Google to find his own old stuff over the years. Some content he was referring regularly disappeared from Google's results. These pages had previously been included in the results.
> Another reasonable method would be to look at aggregate referer trends; is traffic from Google to old pages decreasing faster than traffic from Bing to those pages.
Yes that would be interesting.
I've been wondering whether Google would actually purge the URL too. The Googlebot used to be very persistent in retrying "404 not found" results.