Google emphasizes popularity over accuracy
superhighway98.com
superhighway98.com
Two more examples are error messages and IC part markings --- searches are flooded with results that do not even contain all the words in the query. I didn't put those words there for no reason, ignoring them is absolutely unacceptable. This becomes ridiculous when you search for error numbers, where a search containing the exact number and the word "error" gets flooded with plenty of useless results about other errors.
Currently just to get basically workable results I'm finding myself putting "every" "keyword" "in" "quotes" (to make sure they're actually in the result pages at all), the site: modifier to restrict it to real sites, and negative modifiers (-"keyword") to remove some SEO results.
Google used to be "magic" in that it knew what you were thinking, and gave you what you wanted instead of what you asked for. These days it is just page after page of auto-generated results, pages that don't contain anything relevant to your query, or just low quality results.
I'm not going to pretend I'm an expert in Google's search, and I'm sure they're meeting some metric or another, but from my perspective things have gone seriously downhill to the point where I am looking elsewhere.
It used to be THE technical search engine de jure. Now it feels like a search engine you have to hack to get it to work well. Not a good place to be.
PS - I have read, in HN comments (so pure rumor) from self-proclaimed ex-Googlers, that Google's internal culture punishes people for improvements/maintenance to existing products, and that promotions come from developing new products/features. If even semi-accurate might go a long way to explaining why Google Search feels neglected aside from changes which seem to exist to improve their button line/promote sister products.
But when Google first came out, it was a shock. You could just search for something like "Linux", and the most authoritative sites all showed up on the first page.
At least those search engines gave you that many results to go through... now Google gives you less than that, full of spam (despite the index probably containing far more), and you'll be in CAPTCHA hellban if you try harder to get to the rest.
Of course, you'd have to read a manual to use it and it would have a ton of spam, but some people just want lower-level control - they still sell stick-shift cars.
There’s a long tradition of compiling and publishing concordances, which are just indices of every place each word appears in the original text. They’re generally not useful without access to the original, so noboy seems to mind them very much. Google’s index is just a modern form of the same thing.
On the most extreme end of this, I've seen four-word queries produce results, in which three of the words were stricken out. More often, it's just one word, but it's usually exactly the one that makes the difference between a very specific query, and a very generic one.
Worse yet is that they try to do synonym substitution, but their algorithm has a ridiculously low bar for that. Like, you might be searching for "FreeBSD", and it will substitute that for "Linux", or even "Ubuntu". Or search for a specific firearm model, and it finds "gun".
Quoting keywords suppresses all of that, but synonyms are actually useful - if it did them accurately...
Maybe they're already doing this, but it sure acts like learn-to-rank is always ranking pages as if the query were very sloppy.
It's been a long time, and I certainly never read the code, but I vaguely remember a Google colleague mentioning something (before learn-to-rank) about a back-end query optimizer branch that would intentionally disable much of the query munging if there were any search operators in the query. There was some mention about using cookies / user account information to do the same if the same browser/user had used any search operators in the past N days, but I'm not sure if that was implemented or just being floated as a useful optimization.
Is there a superior alternative to Google?
Cliqz has hardly anything indexed at the moment, but it actually gets relevant results from those. (e.g. "zoom privacy" brings up the Zoom privacy policy first, then three news headlines from the last 12 hours, then a news article from yesterday, then an IT@Cornell guide for making Zoom meetings private, then some more news articles, some stuff about HIPPA…) I really like it, even if it isn't great for programming at the moment.
[DuckDuckGo]: https://duckduckgo.com/ [Cliqz]: https://beta.cliqz.com/
1) Google CANNOT provide you with technical search, because choice of index/query filters is always limited (ie. Do I prefer exact matches over multiple matches?)
2) Google has shareholder & public responsibility. It means that service is adjusted (and it's 'algorithms') towards biggest type of queries performed.
All of this is a constant battle between precision and recall for given query. Adding to complexity, Google needs to account for
* Extraordinary amount of users using their search
* Extraordinary amount of data on webpages
* Importance of authority
In smaller search engines (ie. shop full-text search) you usually adjust towards one use case. This in itself is already hard.
Google does that for all possible use cases, for all possible queries while still fighting same precision/recall battle.
To be clear. I think google is terrible, but I also think that there is no other option for them at this point.
All of this became clear for me the moment I've got interested in build search and relevance engines.
To be fair, in my case the lack of information is real (I've gotten only a minuscule amount of info by asking on niche forums sadly), but cutting out the noise early would be helpful.
Which makes their mission statement "to organize the world's information and make it universally accessible and useful" come across like a bad joke. They really should update it since they don't even bother to keep up appearances anymore.
Thoughtful critique that adds information (like userbinator's GP comment) is of course welcome.
You may be overlooking where Google found the missing words or concepts in a page that linked to the hit, just not within the hit itself. That can actually be useful sometimes.
Then why not just show that as the result.
That can actually be useful sometimes.
...hence why the related: operator exists, it's just not the behaviour that most people expect by default.
The ones that do best, the ones that are the greatest public benefit, seem to be governed by a strong leader who stays at the helm.
As soon as the people at the top change, you get jostling and unclear leadership, and power gets diffused.
You get emergent behavior - lots of internal and external competing for power and interests - and the ability to say "no" or "this sucks" or "we will do this" happens less frequently.
Normally competition would take care of all this, but with big gorillas that dominate a market, it might take a while.
Not necessarily. When you find yourself with beancounters and MBAs at the helm, they for sure will optimize the company to outcompete others. Such companies will eventually die from the rot, but not before they drag down the entire industry they're operating in.
I would agree. The quality of the company seems inversely proportional to the number of employees, https://www.cnbc.com/2020/01/02/google-employee-growth-2001-...
In the early 2000s, it was the one search engine that returned relevant results. All the old ones returned a hodge-podge of key-word matches. Google used the number of citations as a strong signal of relevance (the number of pages that linked to a page meant something).
In the age of GMail beta, Google was my hero and I wanted to work for them. They were doing new things with the browser. First was Google Pack, which included necessary software like Firefox. Then Google introduced Chrome, which was even better.
Both began to descend. I don't know, they became slower and clunkier, as if they were designed by the denizens of any number of companies. Google had about 50,000 people at this point.
Recently, I can't stand even to read Google's documentation, because of its (1) bureaucratic wordiness, and (2) cluttered layout that reminds me of that video about Microsoft redesigning the iPod box. In the early 2000s, Google was a maverick. Nowadays, Google is hard to distinguish from any other corporate giant.
Google is indistinguishable from any other corporate giant because they now are any other corporate giant.
Did anyone seriously think an Ad company would go any other way? It's the scummiest of industries. It was always just a matter of time.
And now we've given (allowed, stood by, whatever) them ALL the data for >50% of our cellular users for the last 10+ years. I'm an Android user too. Oof.
For those interested in the video
SEO spam instead of car repair manuals is a joke. It has a lot bigger and scarier implications. Just try imagining someone like Obama winning elections now... simply not possible... And it's not because people became more dumb, they totally didn't: it's just because now dumb people have a voice online. And spend money online. Internet works for them now, and shows everyone what they want to see...
Cloned a git repo where I thought I remembered seeing the function and 500 MB of downloading later grep confirms that I remembered correctly and that the exact keyword I had been searching for is present multiple times in the source code.
My hunch is that since coders use an ad blocker anyway, it's not financially viable to operate source code search on the public internet.
But I perceive this as Google losing the battle with SEO. On one hand Google writes the algorithms, and on the other SEO tries to exploit these algorithms to rank as high as possible.
I was thinking that a solution would be instead of page rank, have an author rank. This rank keeps track of authors that are experts at certain topics. When you search for certain keywords, the articles of these known experts are ranked higher.
Overnight Oats as an Example. You just find pages upon pages of blogspam. And almost every page has the same useless information. No normal recipe site is somehow in the top 10 pages.
I was searching for the official AWS security certificates (namely, for ISO27001), which AWS neatly publishes on their site. Even for something that specific, the real AWS certification page was at the bottom of the first Google result page. Everything above it was from various random consulting outfits, all trying to sell their "expertise".
When search terms for a company's official security certificate are poisoned with SEO, we know the well has been thoroughly poisoned.
Anyway, I typed in your query " IS027001", and you're right, the AWS result is at the very bottom of the first page on mobile. But if you search "AWS IS027001" it's at the very top.
But IS027001 is not an AWS-specific thing, and the results above it are about the standard. It would seem equally bad to return AWS's product pages to somebody looking for information about the standard.
There's definitely an argument that there are too many results from random websites satisfying the "general information about this standard" intent and it would be good if google could guess what everyone wanted simultaneously, but the query is pretty ambiguous...
I actually do think likes, stars, upvotes, whatever you want to call them are valuable forms of feedback. I like specific reactions - love, laughter, anger - even better. But I do wonder what it would be like if those only went to the poster. If they weren't shown to anyone else, and didn't affect what was shown to anyone else. I suspect that it would make "the feed" a usable model again, instead of the abomination it has become.
The notion that forced filtration, reprioritization, and chronological rearrangement of content is implicit the idea of information feeds (search, social media, etc) shows how deep the brokenness is baked in.
Sometimes the filter is a feature, but sometimes it's not. In the name of convenience, it feels like we've given up our right to choose to see everything and decide for ourselves how to filter.
I think daily about ways to get people to stop donating content to these censorship and surveillance platforms. Most people don't run businesses, so they never realize the rent-seeking nature of these jerks.
This would imply HN "censors" the front page. Of course that's nonsense. Sometimes sorting by an algorithm is just sorting.
In the case of HN it's up votes, and in the case of Google it's a basket of hundreds of criteria (including up votes if you think of clicks on results as voting). Google probably does have some "censorship" rules like filtering out illegal content, but I'd be surprised if they're not impartial about everything else.
That’s precisely what they do. Everyone’s allowed to censor what they wish on their own webpage. I, for example, censor from my own webpage (which otherwise contains a lot of information about me) anything someone could use to physically harm me.
The issue comes up when the censorship is used, for example, in DMs or timeline posts between friends (as it is on Facebook and Instagram and Discord), versus one’s own content on a webpage.
There’s a difference between moderation and rent-seeking.
Assuming impartiality seems awfully optimistic, given both human nature and the fact that Google admits that it is committed to maximizing profit.
When you create the community, it gives you a mad libs style page, where you can choose if you are a democracy, republic, autocracy, a mix of them. Vote weighting, whos votes count for what things, vote prediction and extrapolation. Conditions and Voteing for changing the charter rules, votes for how content gets posted, votes for leadership boards and mastheads. The ability for editors to strengthen or amplify certain voices or voters in a community. If you could seed a community with "role model voters" and then use the other voters who vote like them to extrapolate how they would vote on stories they havent seen. There is this common misunderstanding in the word that votes are only used as votes, and not as a signal to indirectly form a decision.
Every AI feed I've ever used (prismatic being one that comes to mind) buckles under popular vapid content rising above. Extremely strictly moderated feeds do exist tho: https://aldaily.com https://longform.org/ https://longreads.com/ I just wish there was a way for a group of strangers, who arent math experts, could spin up a collaborative feed reader (something that autoingests rss/twitter), and through their collective upvotes posts it to a more static page. Counterparties did something like that with Percolate.com and was able to still use the product once it pivoted markets. https://www.techmeme.com, https://www.mediagazer.com and https://www.memeorandum.com/ do the same thing. But those were all close knit teams that knew each other, their underlying tools arent accessible to spin up your own. I'm sure these tools exist for newsrooms, with the abundance of modern CRMs coming to maturity, a place to chat and edit before things get published. But they arent built for strangers to create community together.
It would make sense for these communities to more resemble Wikis with more static content, instead of the endless feed first. /r/personalfinance wiki being the front page, and the feed being something on the side powering new content to add to it.
Id really like to see something that combines reddit, git, fandom/wikia, techmeme, rtings.com, slant.co, kit.co into a collaborative consumer reports, wirecutter, or metacritic, rotten tomatoes. A place for people to gather to build consensus around something in a more structured way than wikipedia, and then publish it. Places like https://letterboxd.com succeed in some ways, but only through the existence of a shared culture and keeping to a specific topic.
There could also be an element of customizability for the end user. A metacritic or rotten tomatoes where you can weight certain critics votes, or a wirecutter where you can express your preferences for certain traits and then have it output a ranking list. Not completely dissimilar to rtings.com's magic tv ranker.
with the current confinement situation in the world, my podcast episodes list has been growing. i caught up with some of it over last weekend but i feel like i don't have a dedicated time for podcast anymore.
twitter & the web have replaced that time.
It is well worth a read.
Ever since the first time I heard that word, I've always imagined animals on a farm feeding happily at a trough, unaware of their impending doom. It turns out, that's closer to reality than I thought.
Edit: I mean the word "feed".
But probably there is a higher percentage hosted on proprietary platforms (Reddit, FB, ...) that replaces services mentioned above but personal websites as well.
I tried this and there are three news articles about General Motors, an infobox for General Motors, 4 search results for General Motors (their web page, another of their pages, a New York Times article about them, and the Wikipedia article about them), a box containing tweets from General Motors, two more news articles about General motors, and finally a link to GMail. Then the other "GMs" start, including GraphicsMagick for node.js. I think they did a pretty good job interpreting "GM" here, and I don't think the Internet is exactly ruined.
If I type it into the browser search bar however it is the first result. This is not Google's fault, as I'm on Firefox, which is smartly trying to look for sites I frequently visit.
The rest of the results were GM related news/Twitter/links.
For me, a Google search returned only General Motors results in the top spots, while a DuckDuckGo search had Gmail as the first result after the news stories.
"Google sometimes gets things wrong" would be a more accurate title. It wouldn't get any upvotes, nor deserve any (surely it's obvious that a website trying to be as many things as Google does will sometimes be inaccurate). But it would be more truthful.
Example: if(Donald_Trump.as(President) == Bad){ Election.vote = !Donald_Trump; }else{ Election.vote = Donald_Trump; }
I am aware that this comment is more or less useless to many, but I still wanted to write it, because I have the power and the freedom to express my _opinion_, and because I still hope that there are people that will understand what I'm trying to say.
Let Google be Google.
If Google's vision sings the same song as the song of the people, then the people will use it. If it doesn't then people will find another search engine that will sing a song that people will like.
Or in other words, I don't see a reason for us to try to change Google, if Google will want to be the primary search engine, then Google will have to change on it's own.
Am I completly wrong? I Google strong enough to be able to shape the song of creation (strong enough to be able to prevent any other search engine from rising up even though the other search engine could be better)?
This is not what happens easily. In a near monopoly, the monopoly product can be significantly worse than other options, or what is reasonably possible, but people will continue using it for a long time because of various economic, network, psychological lock-in effects.
So it favors popularity over accuracy. So what? Google is not in the business of providing the most accurate search results. It's in the business of generating ad revenue. If popular links make Google more valuable, then favoring them is "getting things right" from the perspective of a private company.
But it's also just a hard problem. That featured snippet about the dentist for example: Google's computers aren't investigative journalists. The purpose of the featured snippet is actually to favor accuracy over popularity by deferring to journalists when Google senses that the searcher is looking for information about an event that was covered in the news. However, if the journalists get it wrong, how is Google going to know? Dollars to donuts, if Google actually knew the right answer, that's what they would surface.
I'm not the one making that claim, I'm just starting from that assumption since the parents are claiming as much
As the defacto front page to the internet, it is aeguably in society's best interest for search results to return more than SEO spam and ads.
I think the modern state of Google is a huge disservice to civilization, compared to what it was and could be. By prioritizing popularity it reduces the majority of search queries to the lowest common denominator and encourages shallow, non-technical culture.
I think what we're seeing in the refinement algorithm is a regression to the layman's mean, so to speak, as they tune (train?) The algorithm to work better for the majority of their users, who happen to be non-technical.
But when you excessively dumb down technology you reduce incentive for people to learn anything and, more importantly here, the dumbing down means showing entertainment and SEO results over possibly more technical content.
Personally I find it disheartening when I search for technical words and the only results are celebrities or media.
Google isn't wrong – it's actually doing its part really well. It's either competition or regulation that need tweaking. Probably both.
In other words
> an economic system that by it's very nature can only provide sub-optimal results is broken.
The economic system can provide optimal results. Individual actors can but don't need to in order for the prior statement to be true.
But this doesn't really apply anymore once service X dominates the way a large proportion of the western population are trying to locate something - as opposed to clicking on links or bookmarks.
Reality is what actually matters. Not what people think or what they wish. It is what people do and what actually happens.
The reality is that Google not only holds a vastly dominant market position in search - but its search is the default way a majority of people interact with the Internet.
The address bar in your browser is not just an address bar. It is both a search bar and an address bar. And that makes a real difference when the search engine behind it is most often one search engine.
But Google's propensity to reward sites/pages that are popular or new rather than those which are actually more accurate/better in terms of quality is definitely an issue.
Well, no, I don't. But I highly suspect that they exist and would want to chat.
Bring back some variety of DMOZ, perhaps in a federated (easy to fork) version. That was quite successful at surfacing the best-quality online resources by topic, and even the early Google index seemed to rely on it quite a bit. But it wasn't a VC-funded project, of course.
1. It was slow to add categories/sites, which especially hurt categories where things change pretty quickly (gaming, tech and media are good examples, since new systems and frameworks need to have categories added ASAP).
2. Editors were often drawn into corruption, and either judged submissions based on how much they were paid elsewhere or prioritised their own/friends/family's websites.
Both of these issues could potentially be fixed with some more resources and better oversight, but it may mean any future DMOZ equivalent would need a lot more funding than the previous one.
Federation would help with both factors, though. A workable "right to fork" is a powerful incentive against corruption. Notably DMOZ was not federated or "forkable" in any real sense, even though it did have a reasonable amount of sites mirroring it.
1) look for references to source materials
2) check references quality - is reference real? does the quoted text match the text from the reference? is it an academic paper published in a journal?
3) authorship quality - what is the academic "impact factor" score for the author?
4) confirmed viewer reviews - subjective review by confirmed users
5) accessibility score - automated user interface usability analysis
High quality data exists, but it's not much of the ad supported web and not much of what users what to read.
First, start with a whitelist. Hand pick high quality publications, and rank them towards the top. This may tilt results back towards institutions, and away from blogs.
Second, punish similarity. If everybody is reposting AP or Reuters without any additional information, consider them a dupe and don't list them. They can run their portals, but they don't need to show up in search.
It's come up multiple times in this thread, car manuals is a good example. They would be better off throwing away every result they have and hand indexing the good information, than what gets returned right now.
Recipes in particular have turned into a giant story about the way grandma used to do it with a picture followed by the same couple variants with different proportions. Pick winners by hand.
Someone has a finance question, just put boggleheads at the top, instead of whichever 59 affiliate credit card sites sprung up.
Need health advice? Put examine.com at the top above WebMD and healthline. Why? Because a human exper compared them and decided examine is a better first result. You could comb through tens of thousands of sites with a team of hundreds of people, something Google easily has at its disposal. What PageRank had, that seems to be missing now, is a seed of "we trust these most" and let the network grow from there. It tried to find expertise, instead of clickability. It was about getting you the best information first.
Correct. I would do that.
.
>They also have their own bias introduced to their content.
Yes. Good.
I'm not necessarily saying they hand pick the best article for every single story. Although techmeme.com and hn do that to an extent, when they notice a better version of an article, they replace the top link with the better version.
I spent a few days thinking about it not so long ago and I have thought of something rarely mentioned. Don't get me wrong, I don't think I have completely solved the problem, just noticed it changes the perspective.
If I remember well, from my user perspective, the biggest change Google introduced was the ranking by page. Yahoo used to rank by site not by page. Maybe going back to a ranking by site would help creating a good index.
A site would be associated to a number of keywords, say 20 and that's it. That would give incentive to pick the keywords you want to rank for carefully and really be an expert about them instead of having SEO experts deciding which keywords they want to rank for this week and write empty TF-IDF optimized blog posts.
This sort of search engine would not give you the answer to everything but it would give back power to the websites. The information retrieval process would then be 2 steps :
- find a good website
- find the information within the website
Now we are at a point where normal website users have very little ways to be high in search results. And people with money can buy it either with very expensive SEO or with expensive Ads.
Sometimes i wonder why the heck should normal website user even try to please Google. As a normal website owner i dont feel Google gives me as much as it expects me to do.
Dont link to this, use AMP, God forbid to Exchange links. There are books about how to please Google. But what is the point? It all comes to who has more money. I dont, so i will never win a good position.
I think we should just forget all these Google rules because they destroyed the Internet how it used to be. Autodiscoverable.
Edit: I think you've actually managed to find a really terrific example of a case where modern ML systems are going to have trouble, because of how that Wikipedia page is worded in relation to the query.
* There isn't a single phrase that answers the query directly, so the ML model would need to make very good use of context (attention), both within multiclause sentences and between paragraphs.
* There are many different numbers on the page, so the model has to determine the right one. It can't just get lucky by guessing here.
* A wrong number (1.7MB) has close proximity to literal keywords in the query (floppy, disk). (The right number [880KB] does too.)
* The model has to understand and properly make use of "most" vs "unusual" in its decision.
Ever heard of PageRank? Google was literally founded on an algorithm that uses the endorsements of "an angry, misinformed mob" to determine importance/relevancy. Obviously this is only one factor in search results (and may not even be used anymore), but this approach is what has made Google successful.
I don't disagree with the general point that the amount of content in today's Web makes the job of a search engine much harder. Perhaps some of Google's techniques lower the result quality for some users, for some queries. That's a much more boring title for a blog post, I guess.
I started grad school at Brown in 1997 and I remember a talk there, by someone from Google or connected to it, about PageRank. PageRank was still new, Google was still in beta. Free swag from dotcoms was literally growing on trees. But I digress.
I remember that the narrative about PageRank at the time was not about popularity, but about expertise. I really remember that the presenter brought up a possible threat to validity—what about gaming the system?—and pointed out that if you wanted to persuade the system that you were an expert on a thing, you could get lots of people who talk about the thing a lot to point a link at your site. BUT, he says: that shows the system works! If you can get that many people who are at least mini-experts on the thing to point to you, the only way to do that would be if you, too, were persuasively an expert on the thing.
(These were... naïve times in many ways.)
Twenty years ago this was highly persuasive. It was an un-game-able system, because "gaming the system" meant the system changed you. I can be pretty cynical about a lot of things, including and especially Google, but I really do think that even Google itself believed this. The company was founded on an algorithm that rewarded expertise.
It turned out that the algorithm also rewarded angry, misinformed mobs, so cmckn isn't exactly wrong here. But I think it's important to be clear that to the extent that Google's algorithms have always done this, they a) weren't meant to, and b) actually didn't in the very earliest days, because the WWW link structure really did, in the 90s, work like they thought it did. (Then the measure became the target, etc etc)
In the old days it was cost prohibitive to setup a massive network of interconnected fake hosts to boost your rank, but these days with cheap VPSes this is entirely possible and probably even profitable. I'm certainly not the only person who has searched for something and found a ton of differently styled blogs all hosting exactly the same content stolen from some legitimate website and clogging up pages of Google results.
It's why buying expired domains and throwing your totally unrelated content on them works great.
It's why subdomain/folder leasing is a thing where affiliate sites will pay "high PR" sites to reverse proxy a subdomain or folder to them and (that's the important part) link to it from their main site. And boy, does that work. The same content that would be > page 100 suddenly is in the top 3.
There are other factors, but they don't matter nearly as much as Google's "we have 200 factors that contribute to the ranking" stuff makes it seem. You can throw the most atrocious low quality content on a site with lots of incoming links and it will rank at the top.
It's not as impossible as you might think. But it's certainly not easy.
[1] http://web.archive.org/web/19990208021747/http://yahoo.com/
My family actually had a book like and I learned about a bunch of the net culture sites through it.
depending on what i am looking for there are a few avenues i would explore if google wasn't available. github, gitlab, & stackoverflow/stackexchange for code related stuff.
wikipedia for general topics. youtube for how-tos. twitter for the news. newsletters and podcasts for links to new articles.
&c.
Still nowadays the best use of Google for me is to find those tematic forums
And I must say search result quality has declined a lot over the years with the past 3 being worse of the sum of the past 20.
must be weird to buy a domain name and still be mostly anonymous and invisible :)
Web crawlers were introduced in 1993. At that point, gopher probably had more content. Remember WAIS? All the libraries liked that protocol. Encarta was still selling like hot cakes.
I could click next as long as i needed. I could refine query to get better results.
But now result list is extremely limited. Refining query gives the same result.
Google was once a search engine that allowed to discover content. Now, it is not.
You could write an article and Google indexed it and showed it if people searched for it. Now it does not work that way. If your audience visits other pages than yours, it will show irrelevant info from these pages rather than perfect match from yours.
And also Patelisms. Once, a short post was enought for Google to index it. Now it has to be essencially a book. It does not need to answer any question, as long as it has a length of a book and thousands of illustrations.
I wished there was a search engine that finds pages matching query, not guessing answers. Giving the freedom to explore rather than giving cheap crappy answers.
I empirically disagree. For me Google often shows small sites with perfect matches before big sites with vague matches and a few of my small sites also rank very well next to giants.
> I wished there was a search engine that finds pages matching query, not guessing answers.
Why not use quotation marks?
Ps. Quotation marks help in some degree. But the response pool is often very small. Also sometimes quotation marks return broader results than expected to
What does this even mean?
The web wasn't "supposed to" be anything. Though I'm not sure what magic search engine OP actually has in mind and how it's supposed to work.
Besides, one of the modern mysteries is that we're in the age of instant information yet you'll notice how many people will write up an entire comment online or bicker IRL instead of doing a cursory search. I don't think it's the internet creating human stupidity / laziness. Unfortunately we had that long before, and search engines simply try to show the best results with minimal context.
Also, I think discussion around tech would be much improved if we tried to come up with a better idea whenever we go through the trouble of complaining about something. Anyone can enumerate why things are suboptimal, and usually when you try to come up with alternatives, you find out it's just trade-offs with no ideal solution.
Trying to pitch an alternative solution (like how a search engine should work) helps drill down into real conversational bedrock that's much more interesting.
Google, Youtube, Reddit, and Facebook all prioritize freshness. Instead of being exposed to things outside our comfort zones, we take solace protected inside filter bubbles.
Instead of the best answer, what usually floats to the top is the most repeated, the most seod, the newest, the most politically correct.
Google's results are considerably worse than they were and part of that is google trying too hard to think what we want instead of guiding us to ask better questions.
Can you explain what this means?
I see people saying this wrt to Duck Duck Go: "No, the results aren't worse, you just have to re-learn how to search with it" and all they really mean is that you need to stack more context into the search input which is strictly worse since you can already do that in Google if you need to put a finer point on your search.
And they mean that since DDG doesn't know "Elm" is a programming language most of the time, by "re-learning search" they seem to mean adding "lang" to the input. Where "re-learn" seems like a romantic way to phrase this obvious problem of the search engine needing more context. In the same way I had to "re-learn how to use a keyboard" when my crappy 2017 Macbook Pro keys started coming off.
How does your complaint differ from this?
Hum, not exactly. Google will ignore extra context that doesn't match your profile. People coped very well when it didn't.
Said commenter, as they type a comment in a marketplace of ideas where people discuss fairly complex ideas.
The internet is a tool, not a solution. I absolutely disagree that it's hard to find for anyone looking for it. The issue is that most people aren't looking for it, and you can't force them.
A lot of people are just looking for entertainment, and that's perfectly fine. They spent their whole day working hard, come home, and now you want to force them to spend their night studying and discovering new ideas? That may be the internet you want but it's not what people want. The internet can be for more than one thing.
This happens quite often on Hacker News. One of the fastest and easiest ways to accumulate karma is to be the first to post something like "$hated_company has always been doing $horrible_thing, they need to change", which most people agree with, every time a popular story about them shows up. (Thankfully, usually only a couple of these appear in most discussions, and usually people don't spam this.)
> Said commenter, as they type a comment in a marketplace of ideas where people discuss fairly complex ideas.
Hacker News can be great for complex discussions, but it's not free of filter bubbles and echo chambers.
Nothing on or off the internet is. That is why it is important to keep an open mind and read widely and voraciously.
Some people are more insightful than most, but its the exact same system as reddit
Your upboats are time sensitive.
When people talk negatively about bubbles, they generally focus on the couple cases where being in a bubble is harmful, such as when it comes to science, politics and so on, but ignore all the other times where a bubble is good.
Do you read politics you agree with or disagree with? I make an effort to click more on things I think I'll disagree with.
There's obviously a level of notoriety beyond just quality. I don't need the same article 3 times under 3 logos.
Hacker News is, in the most charitable interpretation, an act of charity by pg and YC to attempt to provide a forum for a specific subculture. It's obvious from the goal of the project: "things that hackers would find interesting".
But what most people get from the Internet is not a product of someone with effectively unlimited funds (and legendary software engineering ability) building a website to support their community. They get to choose between eyeball-extraction megasites or clunky off-the-shelf vBulletin or Wordpress clones run by underqualified volunteers. It certainly is nice to be among the target audience of Hacker News -- but what if I weren't?
> The issue is that most people aren't looking for [a marketplace of ideas], and you can't force them
No? We can't? Or we shouldn't? Might that be our hyper-individualist culture constraining our imagination of what the right path (survival) might be?
Our evolution within terrestrial physical reality "forced" us to participate in well-calibrated local marketplaces of ideas, and our psychology evolved specifically so that we had a fine-tuned balance of what we subjectively "wanted" and what we found ourselves coming to believe, despite that initial-condition "want" -- dissenting views in a room have both a repulsion but also a very specific gravity -- a closeness that emerges amongst holders of opposing ideas, when these ideas have manifested and are walking around in human bodies within shared meatspace -- we start to empathize with holders of countering views that we're forced to share physical space with. We talk about empathy like it's feelings for the other pieces of meat, but it's perhaps better conceived as a kinship of one tight bundle of ideas for another. It's evolved and it's ancient and it's a very specific foraging strategy for 2D terrestrial creatures finding information/food under those constraints.
And now, we've designed systems that aren't nearly as clever and well-calibrated as our meatspace selves evolved to be. In the purely physical space we evolved for, we had to share space with people we probably disagreed with, and we developed unique tendencies based on the nature of living on a 2D terrestrial plane. Heck, we'd have different psychology favoured if we made it to this level of the great filter, but happened to evolve in the air (3D grid) or within a more one-dimensional environment or some hyperdimensional space.
Speaking of high-dimensional space: enter the internet. Might our prior foraging strategies and adaptations stacked onto our prior foraging strategies... might they fail now? Foraging strategies are informed by the math of the landscape. (search terms: optimal foraging strategies, radius of perception, levy walk, agent-based modeling, in the vein of [1]) Our psychology is tailored to adapting to physical reality on a plane, and the internet might totally fuck that up. (What is an internet bubble? Maybe it's just my stepping out of the 2D terrestrial grid and engaging through a hidden, non-spatial dimension with some foraging target I can sense near me?) It's like all places are piped into one another, outside physical reality. This isn't Kansas. It's the formation of a hyperdimensional object. It's no longer a 2D grid, and our predispositions and adaptations for navigating such a grid might drive us to extinction.
imho, we DO need to consider "forcing" (as a collective) changes in that infrastructure, or else our hitherto evolved psychology might just as well see us destroyed. So in the end, what people "want" is maybe not the highest principle to hold, because what we individually "want" might be tuned for a dying reality, and a prior foraging landscape.
Heh, anyhow... perhaps this is just a mad-cap rant from a technologist and failed biochemist <3
Furthermore, it's so heavily weighted to only show results from sites with lots of SEO juice. I can't find what I'm looking for anymore.
There are some areas where google really excels and that's in technical searches. Those are really well done.
If true, this would be big news, as the "+" qualifier on search terms stopped being obeyed about a decade ago with the launch of... whatever that now-forgotten Facebook competitor Google used to have. The workaround was to use quotes around individual words, which mostly worked, unless it didn't. But the combination of using quotes and "verbatim" (under Tools/All Results") almost always worked.
[1] The way it failed is sort of interesting. I randomly tried "+rhino +cereal +gm", and the first result didn't include the word "rhino", and considered "gm" to be a synonym for General Mills. Quotes around the individual terms seems to work for this query, even without verbatim.
Based on Google’s advanced operator page at one point, the addition of a plus sign to a search termed either prompted the lookup of Google+ pages or a blood type.
'+' was replaced by adding quotes around each search term.
[0] https://googlesystem.blogspot.com/2011/10/googles-plus-opera...
I have been there since the old times of IRC and I have never seen anything promised.
Did anyone really believe this other than middle aged libertarian hackers from the 90s who kept repeating this in prayer like fashion?
The internet is, as the name suggests, a network. Google also is a sort of network, one that is semantically searchable. But Google Search is not a fact checker. It doesn't have an editorial room for every piece of content it exposes, it doesn't have deep knowledge about everything it throws back at us, and it could not have those things at the scale it operates.
That's not Google's fault either. It's not a search engine's job to correct wrong reporting. It's the job of journalists and news outlets to fact-check information, it's the job of citizens to be critical of the information they consume, and it's the job of governments to facilitate that this happens.
It's our fault that our civic institutions have degraded, and not the fault of Google that it spits our own stupidity back at us. It's not Google's fault that we don't like what we see when we look in the mirror and techno utopianism isn't going to save us either.
It does not succeed as well as it did in the past at its current mission. It has regressed.
I’m fairly certain they’re doing the best they can with what’s available to them, but right now they probably get more accurate results by focusing on what is popular than by blindly accepting the contents of the websites.
I vividly remember the promise, and believing in that promise, but there was never a time that the Internet was this. In retrospect I don't see why we should have believed in it in the first place.
It doesn't work.
Wikipedia does a damn fine job in my opinion. You can find decent, democratized knowledge on almost any topic complete with translation into your preferred language (if not on wiki, then through google translate). Or efforts like MIT's open course ware, which hosts lectures and syllabus materials on a variety of topics. There's also arxiv derivatives and a thriving black market (scihub, libgen, etc) of open access to research papers on any topic you'd like. Social media =/ the internet.
No but in all seriousness, the mistake was in relying so heavily on private corporations motivated by profit. An unpopular idea is an unprofitable one. I'm not sure what a non-profit Google would look like... maybe the same... but I feel pretty confident that a non-profit YouTube for example would likely have a totally different recommendation algorithm.
Google may also be lost in an A/B testing catch 22 feedback loop, where what gets clicked gets elevated, and whats elevated gets clicked. The less straightforward, more clickbait cryptic results get clicked to see what they are. The result that spells out all the answers might not get a click at all. Ive noticed in the last year or two wikipedia not even being on the first page of searches it should be the top two or three results for.
There was never any such promise, no contract signed in blood. It was just the wishful thinking of a group of people sharing a particular political persuasion. Just because a monocultural group of pioneers who are all radical thinkers (But not too radical) build something, doesn't mean that the people moving into what they build are interested in sharing, or even tolerating their idiosyncratic culture or politics.
The web is supposed to be a way to route a bit of information from one computer to another. This is so low-level, and so broad, that it proscribes nothing about the outcome.
Imagine in 1995 asking someone whether MSG is harmful or whether razor blades were ever found in Halloween candies. You would hear fairly confident answers of yes to both. Perhaps someone would feel like checking things, and they'd pull out a book of popular facts or an encyclopedia that happily answered yes (or perhaps no) without much discussion, and that would be the end of it. (The Encarta College Dictionary, which awkwardly straddled the transition from books to the internet, had an entirely uncritical definition of "Chinese restaurant syndrome" as a real thing, for instance.)
Go to Google and type in "is msg harmful" and the featured snippet says that while some people have a sensitivity to it, it's definitely not harmful in anywhere near the amounts used in food.
Go to Google and type in "razors in halloween candy" and the top result is a Wikipedia page debunking the urban legends of the '80s. After it is a YouTube video reporting on an actual case of razor blades being found in candy this past Halloween, which Wikipedia hasn't been updated for. In a single Google SERP above the fold, you already have a more accurate synthesis of multiple sources than any single publication has. (And no, I didn't plan this, I didn't even know about the 2019 cases until just now!)
What about political correctness? There's no shortage of content criticizing either the current president or his political opponents past and present. Or, say, search for "puberty blockers" and you get YouTube videos both in favor of them and against them, an handout from Seattle Children's Hospital on how to get them, an article from The Federalist on why they're clearly dangerous, a paywalled article from The Economist discussing it, an article from NBC News saying they've lowered the suicide rate among trans kids... basically any viewpoint you want or don't want, you can find it.
It's all there.
Before social media, people built their own websites just to put their art out there. Now everything runs through a walled garden, and to your point, it's easier to tweet than to go build a website no one will visit.
And yet there is more content on the web, of greater variety and higher quality, than there ever was on that old web. Youtube is a walled garden, but some of what's on there is brilliant, and there's certainly plenty of non-commercial "for the love" content as well. Twitter is, well, Twitter, but accounts like TheSunVanished are publishing ARGs on the platform. Streaming music, video, gaming, all sorts of platforms are full of creative expression despite also being centralized.
Is all of that really unarguably terrible merely because everyone putting their art out there didn't first go through the trouble of manually writing HTML pages on a shared host first?
People used to download MP3s from AOL chat rooms. There were these people that just served up their music libraries using a chat bot for other people to request downloads.
Once you had enough tracks, you could serve up your own bot to let others share in on the fun.
I can't blame progress for killing culture, but I'm glad that back then, something like serving up an MP3 library in an AOL chatroom was a deeply awesome experience.
>Once you had enough tracks, you could serve up your own bot to let others share in on the fun.
And nowadays people torrent. It's not as if piracy and file sharing are dead or anything.
But I think you're really saying the web is terrible for you, because it no longer resembles what you used to enjoy, and because other people use it in ways you don't find interesting. And that's valid for you, but it's a subjective opinion, not objective fact.
Personally, I'm happy giving up downloading random mp3s from the web to have the depth of content and interactivity the modern web offers. But I care more about content than culture.
It might be easier, but it depends on your target audience how effective it is.
Wikipedia is what the web was supposed to be. Prior to that you had Usenet (for conversation) and services like Gopher, Archie, and Veronica (various sorts of databases). The web was meant to be an accessible front end for that that wasn't as headache-inducing as 80x25 terminal.
https://news.ycombinator.com/item?id=22744171
Teams and editorial boards are great at working together to build things, wikipedia, reddit wikis, fandom/wikia. But the tools they have to calculate consensus (not just the most popular answer but the most correct one, given a cultures beliefs) are stone age. Mahalo.com was an attempt at collaborative information sorting, and https://inside.com/ isnt a bad successor, in theory.
I also think the premise that search/feed are the best options is flawed. The directory was, and stills is a great way to drill down into topics. Search should be more encyclopedia/directory like, and less feed like.
And even in the realm of search, some flags or operators to specify if I'm looking for a historical result, a timely news result, a review, a fact, or an opinion could come in handy.
Google would benefit from humans sitting down, searching things and then going "what should the top results actually be, why are we not there" from a much more exerting their bias standpoint. How are they ranking recipes without a kitchen to taste them? I would totally support a Google recipe division taste testing and ranking technique,to actually know what the best pot pie recipe is. It would be less a waste of money than another messenger product every other year.
We have low-quality content generation not because of Google but because of the low cost of publishing (which the author even mentions). It's exactly why we have email spam: it costs nothing to send. To repurpose a Chris Rock bit, "if sending an email cost $5,000 you'd have no more spam email".
What's the alternative being touted here? No Google? Making things harder to find? Seriously?
I'd say a far bigger issue is people sharing content from and to people who think the same, creating these myopic echo chambers of self-reinforcing beliefs.
Google may do a questionable job at filtering out provably false content but people are way worse at that.
Within such a system it's too easy to foster fear, anger and hate and to propagate provably false information. Anti-vaxxers are just one such group that seem to thrive in this informationless world.
There are also other factors. There is a vast decrease of internet's signal-to-noise ratio due to low-cost content you mentioned. For a while we could rely on google to get us through the jungle to useful information. Now every SEO is targeting google search. Most other search engines use google's results, so no help there.
Blaming google doesn't help, their interests are not aligned with user interests. Author is actually criticizing us and our dependence on this one-size-fits-all source of information. The alternative is building and adopting better alternatives to google search.
With the benefit of hindsight, I'd highly recommend it and confirmed my suspicion that my queries weren't the problem, unfortunately.
It also showed queries are usually only advanced topics as one of my other habits (writing down summarized information in my wiki) lets me usually skip searching online for low-hanging fruits or information I already encountered.
Disclosure, this is my side project. It aggregates results from about 30 different vertical search engines based on the query. Google is actually the source for the web results.
Here's an example search for "GM" (saw that discussed earlier in this thread):
Google shapes the Internet and that shape leaves many things in ruins.
Google shapes the Internet by motivating so many people to produce content to make money through advertising. Most people make virtually nothing, a few people make a lot. There is an enormous amount of content on the internet whose driving purpose of making money is secondary to that of sharing something with the world.
Google shapes the Internet through it's algorithm. Ever read or angrily scroll through pages of BS when trying to find a recipe? The only reason anyone does that is for google. My grandmother's recipe for deviled eggs was stored on an index card in a box on the stove. I bet it didn't have 200 bytes of data. A search algorithm can't do much with that so everybody has to add a grand story about their grandmother, her toenails, and how nice a vacation to the Balkans was which nobody ever actually took. It also has to be on top to force people to spend more time on the site, so they're more "engaged".
I just want to know a good amount of time to hard boil eggs in an instant pot, but fuck me for wanting to know the number of minutes. "Organizing information" became being as much of an impediment as possible and directing me to the winner of the SEO race.
Before Internet advertising was so popular people would put things out there much more just because they wanted to, not for some profit motive. Now everybody doing that is doing it on somebody else's platform, making somebody else money, and often tracking everybody who comes past.
Much of the ruins are "caused" by Google, because Google won the race, next in line would have done the same thing. Probably.
It occasionally brings up the question of how to fix it. That probably requires new transport layers, new browsers, and well chosen limitations. Doubtful it would get off the ground. Decentralized solutions tend to get overwhelmed with extremely unsavory things. A new "browser" based on an entirely different stack would be hard to compete, especially if your goal was eliminating ads and tracking and general money-grubbing.
https://en.wikipedia.org/wiki/Henry_Beard
(I have several of his humorous Latin phrasebooks.)
That's like saying Gutenberg ruined books because now I can find books full of BS.
> Instead, Google has become like the pampering robots in WALL-E, giving us what we want at the expense of what we need.
I'm guessing this is referring to "WALL-E shows how the technology will make us lazy and fat" meme. I've watched WALL-E for the first time quite recently, and I notice little details, and I can say that WALL-E is not showing that message. It's actually explained quite explicitly in the movie that the fatness of people is the side effect of prolonged stay in space, and not of their dependence on technology.
The movie is one huge anti-consumerism screed.
2. especially for businesses, it comes off as unprofessional if you are using gmail, yahoo, etc as your contact address
Try searching for 'retirement for French people leaving in the UK'... It is close to impossible to find anything relevant.
In general, I think we would all be better off if phones just supported calling and texting.
“Worse is better” wins again.
According to a Wikipedia citation (The Guardian) this was the original mission statement.[2]
[1] https://about.google [2] https://www.theguardian.com/technology/2014/nov/03/larry-pag...
Disclaimer: I work at Google.
Like people who complain that the latest macs are “terrible”, perhaps I’m simply not google’s audience any more. That doesn’t make google, or me, wrong.
Google didn't ruined the internet. In fact, the internet isn't in ruin. Perhaps the author should reconsider treating Google, or any search engine, as sources of truth.
And if Google ruined the internet, then why give them more power by using Gmail?
I searched for `dentist pulled ex boyfriends teeth`
You do see the excerpt, but right underneath that you see the Snopes link.
I don't think Google should be in the business of debunking articles written years ago. As long as it's relevance algorithms can brings up contrasting sources, in this case the ABC news article and the Snopes stories. It's bad journalism from ABC that they haven't marked that article as redacted even though it's been proven false.
Recently I remember, seeing that the google card UI for the news marked an article as Satire, because in-fact it was a Satire article. I'm not sure if that's because the original article embedded some information that helped Google discover this.
They do a pretty good job at organizing information and making it available.
Google has been a pivotal center of the internet - information has never been easier to find! Ease of access to information does not, however, alleviate the need placed upon the reader to sift out fact from fiction.
As a thought experiment, would our luddite-esque author friend prefer that Google was the arbiter of truth, rather than trends? I certainly would not.
I think it depends what you're looking for. I realize this is obscure, but I can't find any reference online to a big Facebook Platform developers' conference that happened in 2007. It drives me nuts because I was there at Chelsea Piers with 1,000 people, but it's like it never existed.