Google may rank sites for queries that don't appear on the page at all
unlikekinds.com
unlikekinds.com
Or if you do a search for name surname it may return results for "name" with "missing: surname". That is, how on earth if I search about (for example) Peter Petrelli a result about Peter Thiel would be the same thing? That's a full 50% of my search query that Google is discarding...
<what I'm looking up> javascript mdn
give me developer.mozilla.org at the top every time. It's somehow better than site:developer.mozilla.org for some weird reason.
Then you don't even need a search engine. Just type
m your search term here
in the URL bar to search the site directly.EDIT: I'll add for others that this keyword is stored as a bookmark, so you can find it in your bookmarks and edit/delete/etc.
It must be millions of us, but apparently not enough for Google to decide MDN > w3schools.
Tyranny of the majority I guess.
I increasingly suspect that Google is entering the lengthy death-by-design-by-committee phase that will ultimately unseat them as the top player in their space.
Then you've never worked on a search engine. There's a real bias in the comments here because the way technical people search for technical topics is dramatically different from the way the general audience uses a search engine. Most people use natural-language sentence fragments in search. Requiring exact match for every term in a query would vastly reduce usability.
No major search engine has worked that way since the 90s, I don't know why this is suddenly becoming a narrative on this site now.
No, really. Especially if Google is using search queries and clicks to update ML models - non technical people "pollute" the algorithms with poorly formed queries and overwhelming numbers of clicks on non technical articles.
There's probably enough of a niche now for a genuine technical search engine. Which treats keywords like Google used to, maybe with some regex thrown in and what not. None of this full question nonsense.
I personally believe that allowing people to search with full queries has had a negative effect on society - very little critical thought goes into which parts of your question are actually important, and searching is no longer a learnable skill; just ask a literal question and let Google do all the thinking for you.
Or at least, rather than factoring in the expertise of the clicker, factor in the similarity of the clicker to the new searcher, so as to predict the new searcher's behaviour.
We can see the start when HN complained when Google removed + and said use "" instead.
I mentioned this a couple of weeks ago and I can repeat it: if they pay my tickets and a fair price I'll be happy to hold my course "how to continue being best by not nerfing your market leading product".
Personally I gave up last year and I'm now on DDG. Not perfect but less annoying.
DDG has a feedback mechanism (the almost invisible "Send Feedback" thing in the lower right corner), but I'm unsure if anyone there actually sees the feedback and cares enough to do anything about it -- as I said, the problem is getting worse over time.
And while I'm on my soapbox, I'm going to continue complaining about having to use quote pairs instead of + for required terms. That still _really_ grinds my gears.
> Most people use natural-language sentence fragments in search. Requiring exact match for every term in a query would vastly reduce usability.
He wasn't talking about that, and actually said he was fine with synonyms (and I would assume stemming, etc).
His problem was Google flat-out ignoring some of his search terms, thus giving irrelevant results. I don't have examples handy, but I'm under the impression that Google often ignores some of the more relevant terms in my queries to give me more popular results.
Which is ironic since most pre-Google search engines defaulted to OR'ing terms together, and one of Google's innovations was to switch to defaulting to AND. Now they seem to be revisiting the mistake of their competitors and wandering back to something like default-OR.
what is really amazing is how bad youtube search is. I can search for the exact name of a video, and it wont show up in the results. many times I have to use google search to search youtube
How many people have you actually talked to? Over the last few months I've had more than a few non-tech friends/family/etc complain to me about Google's broken search results, often asking if I knew of a way to fix it. The "missing: <query_term>" behavior is particularly enraging: I've seen multiple people yell at Google's search results variations of "F* you, Google! Stop telling me how you're ignoring half of my search!"
> Most people use natural-language sentence fragments in search.
Most people have a wide variety of behaviors that cannot (and shouldn't) be reduced into s single group. That said, a lot of people have learned to use Google in ways that work for them, even if it sometimes seems inefficient or unusual to those of use used to using query languages. Assuming you know what they "meant to ask" is often wrong and usually considered offensive. Priding tools people can choose to use is great; attempting to DWIM[1] - interpreting what the user intended instead of what they said - has been a terrible idea that ruins tools even since it was first attempted >50 years ago at Xerox PARC[2].
I understand the idea, but you are wrong or at least not completely right (I can't speak for all markets).
I know someone who learned to spell as a kid using Google, because if he didn't spell the words correctly he wouldn't get to see pictures of the moon or whatever he wanted to see.
Also after they learned to guess what people meant, Google used to be able to respect +, "" and also the verbatim option hidden beneath a button.
For a while they would also ask politely: did you mean x?
Since then it has become worse and worse: silently rewriting, fuzzing without asking etc.
The problem is that Google can't even get the right search results with natural-language sentences or sentence fragments.
Just today I had to ask Google Scholar for a lot of things since Google itself was completely useless. So it's not like it can't be done. Maybe that's the way going forward, one general search for grandma and specialized search engines for everyone else.
Yes, most search engines have ignored commonly-used words since at least the late 90s. But Google only started doing its idiotic "Missing: [key word from your search]" fairly recently. And it's a step back.
Although I personally don't think this is the case, it's something that's going to provide ammunition to the conspiracy theorists who'll swear they're just doing it to drive clicks to pages with Google ads.
I think a lot of it is about giving too much weight to (its flawed) geolocation.
Last night I searched for "Vintage computer store $location" while I was in $location. All I got back was Best Buy locations, eBay listings, irrelevant Yelp lists, and local newspaper articles about computers in general.
Today I'm in a different place, but performed the identical search, including the $location where I was last night, and guess what -- helpful results!
Google is trying too hard to be Yelp, and not hard enough to be a search engine.
My problem with this is that the usefulness of web searching seems to have been sacrificed in order to be good at answering the common questions. It's great if I have a question a bunch of other people have asked too, not so good if I'm looking for anything more than that.
Course it makes it ever more useless for the difficult, the rare and the old. The personal homepage, blog, or random site with the best knowledge on something esoteric almost never shows any more.
Bing isn't much better, DDG became best - it's certainly not worse any more - almost by default.
In 2006 I would never search "how do I ___", "what is ___", because "how", "what", etc. were just noise, and sites weren't formatted as a question/FAQ like that. I knew to use a series of keywords to find pages that contain the content I was looking for.
I wonder too how much the decline of personal homepages, blogs, random sites, etc. has to do with how much harder it has become to find them in Google.
It really sucks being a frequent traveller, and frequent VPN user and having to REALLY go out of my way just to get some useful default behavior.
VPNing through Iceland and now all the Google Flights prices are priced in Icelandic Kroner and not even a clear dropdown to change this preference. Try the URL hack to change the currency code in the get parameters and MAYBE get what I want.
There are so many ways to detect this and assume this preference, and they opt for the most debilitating one. Many services are following this trend, all because of blindly following some A/B test iteration.
Add the human sense back into it.
And god forbid the VPN server you used was used by someone else doing something odd, welcome to CAPTCHA HELL! Where we make all the assumptions about what you want until search is no longer useful for you, but don't let you access search at all by assuming you the current user are the problem based on IP address alone!
There are more and more reasons to VPN through different countries.
There are more and more teams using A/B tests and completely removing the human element from their UX design process.
A human would say "hey it makes sense to have a currency picker here, thats super easy since we already have this on the backend"
The A/B test says "hey nobody taps the currency picker so therefore it must be REDUCING your engagement" When its really completely unrelated.
I wonder if NavigatorCurrency.currency should be a thing!
This didn't used to be the case. It used to be, if I searched for it, I could find it, no matter how popular or controversial it was. I used to live in China and it reminds me of how Baidu works.
Citation? Examples?
That's a really hard issue to solve...
I find the message "missing: [term]" more helpful than not. Anytime I see that, I don't click the link. That sends a small signal to Google (compounding with the scale of search queries) that their #1 link isn't what the user was looking for.
A hilarious example of this last week; searching 'email CTR' brought up pages with email/center... Google of all companies should care about CTR!
https://en.wikipedia.org/wiki/Google_bomb
"Google bombs date back as far as 1999, when a search for "more evil than Satan himself" resulted in the Microsoft homepage as the top result.
In September 2000 the first Google bomb with a verifiable creator was created by Hugedisk Men's Magazine, a now-defunct online humor magazine, when it linked the text "dumb motherfucker" to a site selling George W. Bush-related merchandise."
I speak two languages, english and german. I set my search preferences accordingly and expect the results within those languages to be ranked by their actual quality. However it seems Google favors results in my local (german) language, which often don't appear to be the best results if both languages were weighted equally.
Particularly when searching for "{popular game/album/film} review" I perceive the returned results as subpar, simply because my local media landscape doesn't seem as thorough, high-quality and in-depth to me.
Personally I use Google only for local searches and DuckDuckGo for almost everything else. Works well for me.
For basic keyword searches. It’s not uncommon for search users to want inexact equivalencies, synonyms or conceptually close search results. (Search for cold, get back rhinovirus). Every user and use case has different definitions/tolerances of “this term appeared on the page”
It’s a complicated topic, without simple solutions. I wrote about it here.
https://opensourceconnections.com/blog/2018/12/07/synonyms-b...
We can argue whether there is place for more operators, or whether it's better on average for an average user to still be "outsmarted" by the engine, but it's clear what's going on - these aren't exact match results and the results often have zero relevance with what user is looking for precisely because the user knows the specific word WILL appear on the page they need and a page without that word WON'T be what they need. I don't know how much clearer we can be :) It used to work (or seemed to). Now it doesn't. It frustrates some people. That's it.
I've interpreted searches. Database searches, not web pages, but my problems should apply to web search too: For example, a user who types "märz 2019" with quotes might mean that string, in German, or might on the other hand mean that particular month and use the quotes to eliminate february 2019, march 2018, etc. And people enter the exact string they remember in the hope of avoiding a sea of mismatches, but then they either mistype or don't remember quite the right wording.
Don't mess with my search terms! Bring exact matches back! Only show pages that match the search terms!
The amount of automatically created synonyms is getting out of hand, it decided the name of where I work is a synonym for another company in the same business that's 15 miles away
It seems to me the main failing of Google is they have no good way to say "We understood your question, but the answer isn't on the internet".
Instead they just return a set of not very relevant results.
This has been abused many times by using derogatory link texts when linking to politicians web pages.
https://github.com/HTTPArchive/legacy.httparchive.org/blob/m...
You can search for anything in the source-code of pages, using any regex or grep you can imagine.
Obviously running your custom filter logic across every byte of data that has ever existed on the internet is compute heavy... But Google has lots of that!
Unfortunately, it is no longer possible to use Google Search for searching in Internet. You are only allowed to look up viral memes and query neural networks, trained by other people's searches. Searching for rare and unique things will quickly get you banned. An interesting side-effect: using modern Google to check if name is vacant is meaningless. I used to google for names to see, if something else used them, but Google's search repeatedly returned me 0 results, even when there were multiple pages with that particular name in title, many of them years old. Searching for the same thing couple of weeks later have suddenly returned hundreds of hits.
I wonder, how much it costs to actually query Google's database instead of some distant neural network approximation. Apparently, Google's own employees can't afford it anymore — some of them are using DuckDuckGo instead (they even added it to Chrome, lol)
And if I were to hazard a guess I think you dramatically overestimate the role of neural networks in search in general.
Then consider yourself lucky. The post was a little rambling, but it wasn't wrong.
This happens often when creating complex search terms with 'inurl' and 'intext' operators. This happens to me at least once a month.
I suppose it sometimes looks like my account is used for botting, but then again: why wouldn't I be allowed to connect my bot to their results, given that google uses bots to index the content in the first place?
Despite the fact that I could probably adjust to better fit their 'nothing weird going on here' preconceptions, I think I'd rather complain and hope that over time it's google that will change.
If neighbouring IP's to yours abuse Google services, you'll see way more of these errors.
You get warnings and captchas.
Frequently they are unsolvable captchas that just waste your time. :(
"azure" devops agent "checkout" "freeze"
I'm obviously searching for something very specific here, and Google helpfully decided to show me results omitting "freeze" and "checkout" so I had to quote them.
I also notice that it looks like Google thinks that check-up is the same thing as checkout, which in context are not synonyms, though this would require Google to infer "git checkout" but I didn't include git because that would introduce another universe of unhelpful results.
https://dzone.com/articles/getting-started-with-jenkins-the-... was included which doesn't include "azure" at all, though this is obviously Google being cheeky.
Of course I could also call this somewhat sinister, as Google is basically saying I should ditch a direct competitor to their services and instead use a self-hosted service. Or maybe I'm projecting since I would like to ditch Azure DevOps and use Jenkins, but in any case there's an example.
Feel free to play with it: https://www.google.com/search?rlz=1C1CHBF_enUS820US820&biw=1...
Of course, in this case it was a networking problem so I do have to admit that usually when Google starts ignoring what I typed it's because the answer to the question is something else entirely. Though sometimes I know exactly what I'm asking for and Google doesn't get it.
I think possibly what is happening here is that the index believes "azure" is on the page.
Consider the second result here: https://www.google.com/search?q=%22azure%22+%22Allows+adding...
(the one that starts java.dzone.com but leads to the same article).
It appears to have captured at index time a phrase "Microservices and Serverless on Azure" that is no longer in the artcile (possibly a link to another article).
Perhaps a Googler can explain exactly what is happening.
Because anyone who's used Google in the last 10 years knows that Google tailors its results to what it thinks you want to see.
So a search query that fails for one HN reader may work for another, or at least work differently.
A repeated frustration in my life has been explaining to bosses that the search results that appear on their personal phone are not the same as what will appear on a client's computer.
Do you have an example? Exact match has been working fine for me.
Books:
Lucene in Action http://manning.com/books/lucene-in-action Solr in Action https://www.manning.com/books/solr-in-action Relevant Search http://manning.com/books/relevant-search (disclaimer, my book) Elasticsearch Definitive Guide (free/online) https://www.elastic.co/guide/en/elasticsearch/guide/current/... Introduction to Info. Retrieval https://nlp.stanford.edu/IR-book/information-retrieval-book.... Deep Learning for Search http://manning.com/books/deep-learning-for-search
Training:
Lucidworks Training https://lucidworks.com/resources/solr-training-and-consultin... Elastic's training https://training.elastic.co/ OpenSource Connection's training (disclaimer, my company's training) http://o19s.com/events/training
Conferences:
Haystack http://haystackconf.com (disclaimer I'm a co-organizer) Activate http://activate.conf SIGIR http://sigir.org ECIR http://ecir2019.org/ Search Solutions https://irsg.bcs.org/SearchSolutions/2018/ss2018tutorials.ph... Berlin Buzzwords https://berlinbuzzwords.de/ MICES E-Commerce Search - http://mices.co
Blogs:
OpenSource Connections blog - http://o19s.com/blog (disclaimer - my companys blog) Lucidworks Blog - https://lucidworks.com/blog/ Sematext Blog - https://sematext.com/blog/ Elastic's blog http://elastic.co/blog
I'd hazard a random guess that 90% of the time someone looks up a phone number for an unlisted number, they want to know if it's a scammer/spammer. So why the author seems to think this is the "wrong" result to show seems pretty strange to me.
And the algorithm in general appears to frustrate many people, if the comments here are anything to go by.
https://searchenginewatch.com/sew/news/2126346/google-introd...
Verbatim search is still available today. Searches from the URL bar can be made to use verbatim search by default via this search string:
"Stars! game" (with or without quotes) helps a little. The right wikipedia page is linked. Mostly it's drowned out in a sea of stuff about the Dallas Stars, NBA all stars, etc.
"Stars! game windows" is getting closer and has more results, but it's still cluttered up with results for various windows games, companies (StarDock), etc. that have nothing to do with what I want.
The real trick here is to search for Stars AutoHost, since that was one of the hub websites of the game. Doesn't need to be verbatim, quoted, anything. Suddenly, the actual page titles actually have "Stars!" - complete with exclamation mark - in them.
This isn't an edge case for me. C++, C#, ".NET", tons of tech thinks punctuation is cool and (ab)uses it. To say nothing of all the various operators these languages have. Bleh.
If I am searching for "xyz abc" I am searching for those specific terms for a reason, so please don't present me with results for the entirely unrelated "123", and no I am not talking about showing synonyms.
Posted about it at length here: https://neosmart.net/blog/2016/on-the-growing-intentional-us...
Link is from 2016 obviously.
Lately I'm using more natural language searches, because they return better results. Probably they're optimizing the engine to return the best results for the most people, and most people want to know "how do I turn on the furnace" not "[model number] furnace manual"
This happens regularly when a web page is blocked from crawling via Robots.txt file. Google still indexes and ranks the page, but Google has no idea what is on the page.
Want to keep a web page out of rankings? Allow Google to crawl it (by not restricting via robots.txt), but use a NoIndex meta tag or X-Robots-Tag HTTP header to indicate that the page should not appear in search results.
Fix the problem with X-robots-tag, request removal from index (*while such option is available) and then wait for months for the penalty to be lifted.
https://en.wikipedia.org/wiki/Google_bomb
> Google bombs date back as far as 1999, when a search for "more evil than Satan himself" resulted in the Microsoft homepage as the top result.
Basically, Google tries to anticipate query refinements, improving the search as a skilled user would do. Most users are not skilled though, so this leads to an overall better search experience for them, but makes the experience for skilled users worse, due to false positives.
Or - way simpler - maybe a link to the page at the text "4843218317" on it.
So Google has clearly determined that this number is somehow associated with sketchy Uber SMS messages, out of a very sparse signal...
That seems pretty cool actually. Like, really cool.
If there are 10 million websites out there linking to a single article when referencing X, even if such article doesn't include X verbatim, that article will rank high in queries for X.
Includes the unwanted page: https://www.google.com/search?q=2109085405
Does not include the unwanted page: https://www.google.com/search?q="2109085405"
Maybe not the most compelling example, but I think it's usually with really blatant misspellings that Google will second guess my quoted terms, which can be annoying if I'm actually looking for the misspelled term.
- The exact search for that string returns nothing
- You can still force it to do that quite easily by clicking "Search instead" right at the top
If you search for a term with no occurrences which is very likely to be a misspelling of an existing popular result, what maximizes total usefulness for users?
- No results (10% of the cases you are actually searching for something that does not exist)
- Here are the results for the term you're likely searching for (90% of the time) + btw click this link right here to get (no) results for the nonsensical search term
What surprised me most is that there actually exist pages containing both of those terms! But it's also including results with just one or the other, and at least one page with neither term in it. I'm also getting no "Search instead" or "Did you mean" links.
The problem with the "barchart.com" results is that it's some sort of a portal of frequently updating headlines, so it's not reasonable to expect the search index to always keep up to date with it.
I rest my case :)